<?xml version="1.0" encoding="utf-8"?>
<?xml-stylesheet type="text/xsl" href="/feeds/rss-style.xsl"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>Aman Deep</title>
        <link>https://leamandeep.github.io</link>
        <description>Bridging machine learning and Earth observation. I build AI systems that analyze satellite imagery and geospatial data to uncover environmental insights, track climate trends, and support sustainable development.</description>
        <lastBuildDate>Thu, 03 Sep 2026 05:53:13 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>Astro Chiri Feed Generator</generator>
        <language>en-US</language>
        <copyright>Copyright © 2026 Aman Deep</copyright>
        <atom:link href="https://leamandeep.github.io/rss.xml" rel="self" type="application/rss+xml"/>
        <item>
            <title><![CDATA[Language Meets Remote Sensing: Applying VLM to Land Cover Segmentation]]></title>
            <link>https://leamandeep.github.io/language-meets-remote-sensing-applying-vlm-to-land-cover-segmentation</link>
            <guid isPermaLink="false">https://leamandeep.github.io/language-meets-remote-sensing-applying-vlm-to-land-cover-segmentation</guid>
            <pubDate>Wed, 02 Sep 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Introduction In this Blog, I will discuss the LSeg (Language-driven Semantic Segmentation) paper. In recent years, semantic segmentation has gained significant traction in computer vision. Incorporati...]]></description>
            <content:encoded><![CDATA[<h1>Introduction</h1>
<p><em>In</em> this Blog, I will discuss the LSeg (Language-driven Semantic Segmentation) paper. In recent years, semantic segmentation has gained significant traction in computer vision. Incorporating natural language adds another layer of complexity—especially when asking a model to predict unseen classes, a key limitation of traditional architectures. LSeg addresses these exact challenges.</p>
<p>While few-shot learning approaches exist, they still require labeled training samples for new classes. In contrast, the authors focus on zero-shot learning. By leveraging language embeddings, LSeg can identify both seen and unseen classes without additional labeled data. Inspired by CLIP’s strong zero-shot capabilities, the authors adapt this concept to dense prediction tasks.</p>
<h1>Main Idea</h1>
<p>The authors’ core approach is to map both image pixels and text labels into a shared embedding space, allowing the model to assign a semantic label to every single pixel.</p>
<h4><strong>Text Encoder</strong>:</h4>
<p>The text encoder embeds $N$ potential labels into a continuous vector space $R^C$. To accomplish this, the authors utilize CLIP’s pretrained text encoder.</p>
<p></p>
<h4><strong>Image Encoder:</strong></h4>
<p>Similar to the text encoder, the image encoder generates a dense embedding vector for every pixel. The authors use Dense Prediction Transformers (DPT) for this component. Given an image of size $H \times W$ and a downsampling factor $s$ <br />
 the model outputs dense feature embeddings of shape   </p>
<p><img src="../../assets/images/posts/language-meets-remote-sensing-applying-vlm-to-land-cover-segmentation/Untitled-2026-09-01.png" alt="Lseg's architecture" /></p>
<h4><strong>Word-Pixel Correlation Tensor:</strong></h4>
<p>Once the image and text features are embedded, the model computes their correlation using a dot product (or cosine similarity). It will create tensor of size  defined as </p>
<p>&lt;Math
expression="f_{ijk} = I_{ij} \cdot T_k
\[25pt]
where ~ I_{ij} \in R^{\tilde{H} \times \tilde{W} \times C}</p>
<pre><code>T_k \in R^{C \times N}
"
/&gt;

For  N = 1, embedding size of Image $$I\_ij $$ and Text $T\_K$

&lt;Math
  expression=" T_k \in R^C 
\\
I_{ij} \in R^C
"
/&gt;

Below is the pixelwise softmax

&lt;Math
  expression="% Requires \usepackage{amsmath} in your preamble for \text{}
y_{ij} = \text{softmax}\left(\frac{f_{ij}}{t}\right) = \frac{\exp\left(\frac{f_{ij}}{t}\right)}{\sum_{i=1}^{H} \sum_{j=1}^{W} \exp\left(\frac{f_{ij}}{t}\right)}

\\[10pt]

t: Temperature
"
/&gt;

During training, LSeg minimizes the standard spatial cross-entropy loss across all pixels, encouraging the model to align visual pixel representations directly with their corresponding text label embeddings.

#### Spatial Regularizer Block:

To address memory constraints during training and inference, the authors introduce two spatial-text fusion blocks:

* **DepthwiseBlock:** Applies depthwise convolutions followed by a non-linear activation to efficiently fuse text and image representations.
* **BottleneckBlock:** Further enhances the depthwise convolutions by incorporating the output of a max-pooling operation over the set of text labels, capturing the most prominent label interactions.

Finally, a bilinear upsampling block restores the spatial feature maps from &lt;InlineMath expression="\frac{H}{s} \times \frac{W}{s}" /&gt; back to the original image dimensions ($H \times W$), yielding high-resolution pixel-level predictions.

# My Implementation

Unlike standard zero-shot setups, my objective was supervised remote sensing segmentation using a lightweight foundation model (`prithvi_eo_v2_tiny_tl`) aligned with a pretrained CLIP text encoder.

**Methodology Highlights:**

* **Decoder Modification:** Kept Prithvi's spatial decoder, removed the default classification head, and added a projection layer to map dense spatial features into the text embedding space. An encoder-only baseline yielded low-resolution, "blobby" segmentations; retaining the decoder restored finer spatial boundary definition.
* **Pixel-Text Alignment:** Evaluated dot-product similarity between projected spatial features and text embeddings, applying a pixel-wise softmax to derive output class probabilities.
* **Fine-Tuning:** Trained for \~12 epochs using a 300-image subset of the Prithvi dataset.
* **Prompt Engineering:** Applied mean centering to neutralize class biases and prompt templating (e.g., *"remote sensing view of \[class]"*) to enrich text feature representations.

### Key Improvements Made

![Pre-mean-centering similarity matrix for prompt vocabulary vs dataset classes](../../assets/images/posts/language-meets-remote-sensing-applying-vlm-to-land-cover-segmentation/Untitled-2026-09-02-1655.svg)

**Heatmap matrix showing pre-mean-centering cosine similarity scores between fixed dataset target classes (x-axis) and prompt vocabulary terms (y-axis).**

![Post-mean-centering similarity matrix for prompt vocabulary vs dataset classes](../../assets/images/posts/language-meets-remote-sensing-applying-vlm-to-land-cover-segmentation/Untitled-2026-09-02-1656.svg)

**Heatmap matrix showing post-mean-centering cosine similarity scores between fixed dataset target classes (x-axis) and prompt vocabulary terms (y-axis).**

# Limitation

While LSeg provides an elegant baseline for language-guided segmentation, several practical limitations emerged during experimentation on remote sensing data:

1. **Prompt Complexity Ceiling:** The text-image correlation mechanism degrades when exposed to lengthy or multi-attribute text prompts. While basic action prompts (e.g., *"detect water"*) yield clear similarity maps, long-form natural language prompts degrade segmentation precision.
2. **Domain-Specific Vocabulary Drift:** The model's performance drops when encountering out-of-vocabulary spatial features. Satellite imagery presents intricate spatial textures, non-standard orientations, and multi-spectral signals that are far more challenging to align with text embeddings than standard RGB natural images.

&gt; 🔗 Check [demo app](https://vlmsatellitesegmentation-hktm3u6yygaeh4hohbjdnb.streamlit.app/)  and [code](https://github.com/leamandeep/vlm_satellite_segmentation)</code></pre>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Prithvi-EO 2.0 : A versatile Multi Temporal Earth Geospatial Foundation model]]></title>
            <link>https://leamandeep.github.io/prithvi-eo-2-0-a-versatile-multi-temporal-earth-geospatial-foundation-model</link>
            <guid isPermaLink="false">https://leamandeep.github.io/prithvi-eo-2-0-a-versatile-multi-temporal-earth-geospatial-foundation-model</guid>
            <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Introduction In this blog, I will discuss the research paper Prithvi-EO 2.0: A Versatile Multi-Temporal Earth Geospatial Foundation Model. In recent years, a lot of geospatial foundation models have b...]]></description>
            <content:encoded><![CDATA[<h1>Introduction</h1>
<p><em>In</em> this blog, I will discuss the research paper <strong>Prithvi-EO 2.0: A Versatile Multi-Temporal Earth Geospatial Foundation Model</strong>. In recent years, a lot of geospatial foundation models have been introduced in EO, claiming to solve various EO problems. A foundation model is a general-purpose model that is trained on massive datasets and can be fine-tuned for many downstream tasks.</p>
<p>In this paper, the authors mention three limitations of available models. First, these models are either good at spatial or temporal learning. Most of these models are not capable of handling both modalities at the same time. If they are good at capturing long-range dependencies, they may not be able to handle larger areas, and vice versa. EO observation data are spatiotemporal in nature. We need both aspects of the data to understand changes.</p>
<p>Second, most of the evaluations done on these models were not performed on diverse tasks, so clear comparisons were limited. Third, adapting GFMs for EO tasks is challenging and may require a certain level of AI expertise.</p>
<h1>Main Idea</h1>
<p><em>In</em> Prithvi, the researchers used a ViT model. They replaced the 2D patch embedding and 2D positional encoding with 3D patch embedding and 1D positional encoding (sin/cosine encoding). The 3D input consists of (t), (h), and (w) dimensions. The 1D positional encoding is defined as:</p>


<p><img src="../../assets/images/posts/prithvi-eo-2-0-a-versatile-multi-temporal-earth-geospatial-foundation-model/Untitled-2026-03-20-patch.svg" alt="Sin cos embedding in spatio temporal dataset" /></p>
<p>They also encoded metadata such as latitude, longitude, and the day of the year, and added it to the embedding token using a weighted sum of learned weights. Since latitude, longitude, and day-of-the-year information can be missing from the data, the authors simulate such situations by randomly dropping either temporal or geolocation information while masking image patches. Attention is applied across both spatial and temporal dimensions. So, it is able to handle both modalities at different scales. </p>
<p><img src="../../assets/images/posts/prithvi-eo-2-0-a-versatile-multi-temporal-earth-geospatial-foundation-model/Untitled-2026-03-20-1655.svg" alt="Prithvi E0 2.0 architecture " /></p>
<p>Prithvi was scaled to 600M parameters, placing it among the largest models for EO tasks. MAE is used for pretraining. Prithvi was trained on a 4.2 million global time-series dataset.</p>
<p>Prithvi used a stratified and diversity-oriented sampling strategy to create a globally representative training dataset. The authors selected HLS tiles based on <strong>land-use/land-cover (LULC) diversity, ecoregions, and spatial heterogeneity</strong>. They sampled tiles representing different LULC classes, specifically increasing the sampling of <strong>urban areas</strong>, and included tiles with high LULC entropy to capture heterogeneous landscapes. They also ensured broad <strong>ecoregional coverage</strong>.</p>
<p>For each selected location, they constructed <strong>four-image temporal sequences</strong>, with observations separated by approximately <strong>1–6 months</strong>, allowing the model to learn seasonal and temporal changes. These sequences were divided into <strong>256 × 256 pixel patches</strong>. Cloudy and poor-quality samples were filtered out, while spatial overrepresentation was controlled. This produced about <strong>4.2 million training samples</strong> and <strong>45,568 validation samples</strong>.</p>
<h1>Limitation</h1>
<p>The authors mentioned that as the number of timestamps increases, the computational complexity also increases. The attention mechanism used in Prithvi has quadratic complexity by nature.</p>
<p>Satellite images are not continuous by nature. We also have the issue of cloud pixels. I feel this limitation is present in the model, but it is not solving the missing-frames problem. I did not find a detailed discussion of this issue in the paper.</p>
]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Welcome  | Computer Vision & Geospatial Data Science]]></title>
            <link>https://leamandeep.github.io/welcome-computer-vision-and-geospatial-data-science</link>
            <guid isPermaLink="false">https://leamandeep.github.io/welcome-computer-vision-and-geospatial-data-science</guid>
            <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Welcome! This blog serves as an open notebook for my work and exploration across computer vision, geospatial analysis, and applied machine learning. Why This Blog? Geospatial AI moves fast. Between pa...]]></description>
            <content:encoded><![CDATA[<p>Welcome! This blog serves as an open notebook for my work and exploration across computer vision, geospatial analysis, and applied machine learning.</p>
<h1>Why This Blog?</h1>
<p>Geospatial AI moves fast. Between paper releases, model iterations, and pipeline developments, sharing knowledge helps bridge the gap between academic theory and practical implementation.</p>
<h1>What to Expect:</h1>
<ul>
<li><strong>Research Summaries:</strong> Clear breakdowns of impactful papers in spatial AI and computer vision.</li>
<li><strong>Model Highlights:</strong> Behind-the-scenes looks at my custom models, experiments, and code architectures.</li>
<li><strong>Practical Guides:</strong> Workflow breakdowns for handling complex spatial datasets with modern CV tools.</li>
</ul>
]]></content:encoded>
        </item>
    </channel>
</rss>