In August 2018, Kerala experienced its worst flooding in nearly a century after extreme monsoon rainfall. Alappuzha — a low-lying district criss-crossed by the Pamba and Achankovil rivers and the Vembanad backwaters — was among the worst affected. This post is the complete, cell-by-cell companion to the Flood Mapping with Machine Learning & Remote Sensing workshop notebook — every section of the notebook, in order, with the full code, what each line does, and the same discussion questions used in the live workshop. Run it in Google Colab alongside this post, or use this as a standalone reference afterward.
By the end, you’ll understand why clouds make optical satellites unreliable during a flood, how SAR gets around that, how a Random Forest classifier turns labeled example pixels into a wall-to-wall flood map, how to evaluate that classifier honestly, and how to isolate genuinely new flood water from rivers and lakes that are wet year-round.
Section 1 — Setup
Earth Engine needs two Python libraries beyond what Colab ships with: geemap (interactive maps for Earth Engine) and the earthengine-api itself. Run this once per Colab session.
!pip install -q geemap earthengine-api
Now import everything we’ll use throughout the notebook:
import ee
import geemap
import geemap.colormaps as cm
import ipywidgets as widgets
from IPython.display import display
import pandas as pd
import matplotlib.pyplot as plt
Authenticate and connect to your Earth Engine account. The first time you run this in a fresh Colab session, it opens a login flow in a popup/new tab. Before running: replace 'your-gee-project-id' with your own Google Cloud project ID that has the Earth Engine API enabled (required for all Earth Engine accounts — create one free at console.cloud.google.com).
PROJECT_ID = 'your-gee-project-id' # <-- CHANGE ME
try:
ee.Initialize(project=PROJECT_ID)
except Exception:
ee.Authenticate()
ee.Initialize(project=PROJECT_ID)
print('Earth Engine connected successfully.')
Output
Earth Engine connected successfully.
Section 2 — Study Area: A Real Boundary, Not a Box
Instead of drawing an arbitrary rectangle, we'll use the actual administrative boundary of Alappuzha district, taken from the FAO GAUL dataset that's already loaded in Earth Engine's public data catalog. This matters because a rectangle usually includes irrelevant land outside your area of interest (or cuts off part of it) — a real boundary keeps every computation faithful to the place you actually care about.
Load the global district-boundary dataset:
districts = ee.FeatureCollection('FAO/GAUL/2015/level2')
Filter it down to just the one district we want: Alappuzha, Kerala.
alappuzha_fc = districts.filter(ee.Filter.eq('ADM2_NAME', 'Alappuzha'))
Pull out the geometry (the polygon boundary itself) — this aoi variable is what we'll clip every image to for the rest of the notebook.
aoi = alappuzha_fc.geometry()
Sanity-check: how big is this district, in km²?
area_km2 = aoi.area().divide(1e6).getInfo()
print(f'Alappuzha district area: {area_km2:,.1f} km2')
Output
Alappuzha district area: 1,316.7 km2
- Alappuzha is roughly 1,300 km² — compare that to the size of a city or region you know.
- Why might clipping to the real district boundary (instead of a rectangle around it) matter for the area statistics we'll compute later?
Now let's actually see this boundary on a map, so we have a visual reference for everything that follows.
boundary_map = geemap.Map(center=[9.25, 76.55], zoom=10, min_zoom=9, max_zoom=11)
boundary_map.addLayer(alappuzha_fc.style(color='red', fillColor='00000000', width=2), {}, 'Alappuzha district boundary')
boundary_map.addLayerControl()
boundary_map
- Pan/zoom the map. Notice the irregular, elongated coastal shape — very different from a rectangle.
- This shape is exactly what every image clip and area sum below will respect.
Section 3 — Define the Pre-Flood and Flood Date Windows
To detect a flood we need two snapshots in time: one from before the event (the "normal" baseline) and one during/just after the event (when water has spread). We'll compare everything against this baseline. The 2018 Kerala floods peaked in the third week of August. We'll use a dry pre-monsoon-peak window in July as baseline, and a window spanning the flood peak in mid-to-late August.
PRE_START, PRE_END = '2018-07-01', '2018-07-31' # baseline / "normal" conditions
POST_START, POST_END = '2018-08-16', '2018-08-23' # flood peak
- Why use a window of dates (several weeks, or over a week) rather than a single date? (Hint: think about satellite revisit time, and cloud cover.)
Section 4 — Multispectral (Sentinel-2): Pre-Flood Image
Sentinel-2 captures visible, near-infrared and shortwave-infrared light — like a very detailed, multi-band camera. First we need a cloud mask, because Sentinel-2's QA60 band flags pixels contaminated by cloud or cirrus, and we don't want those polluting our analysis.
Define a reusable cloud-masking function. Bits 10 and 11 of the QA60 band mark cloud and cirrus pixels respectively.
def mask_s2_clouds(image):
qa = image.select('QA60')
cloud_bit_mask = 1 << 10
cirrus_bit_mask = 1 << 11
clear_sky = qa.bitwiseAnd(cloud_bit_mask).eq(0).And(qa.bitwiseAnd(cirrus_bit_mask).eq(0))
return image.updateMask(clear_sky).divide(10000).copyProperties(image, ['system:time_start'])
Load the Sentinel-2 image collection, restricted to our district, and apply the cloud mask to every image.
s2_collection = (ee.ImageCollection('COPERNICUS/S2_HARMONIZED')
.filterBounds(aoi)
.map(mask_s2_clouds))
Build the pre-flood composite: take the per-pixel median of every cloud-masked image in the pre-flood window, then clip to our district boundary.
s2_pre = s2_collection.filterDate(PRE_START, PRE_END).median().clip(aoi)
Visualize it as a true-color (RGB) image.
vis_rgb = {'bands': ['B4', 'B3', 'B2'], 'min': 0, 'max': 0.5}
map_s2_pre = geemap.Map(center=[9.25, 76.55], zoom=10)
map_s2_pre.addLayer(s2_pre, vis_rgb, 'Sentinel-2 Pre-flood (RGB)')
map_s2_pre.addLayerControl()
map_s2_pre
- This is the district under normal (pre-monsoon-peak) conditions. Note where the rivers, backwaters, and Vembanad Lake already show up as water — we'll need to remember this later when separating new flood water from water that was already there.
Section 5 — Multispectral (Sentinel-2): Flood-Period Image
Build the same kind of composite, but for the flood date window.
s2_post = s2_collection.filterDate(POST_START, POST_END).median().clip(aoi)
Visualize it the same way.
map_s2_post = geemap.Map(center=[9.25, 76.55], zoom=10, min_zoom=9, max_zoom=11)
map_s2_post.addLayer(s2_post, vis_rgb, 'Sentinel-2 Flood period (RGB)')
map_s2_post.addLayerControl()
map_s2_post
- How much of the district is obscured by clouds in this image, compared to the pre-flood one?
- If you were a disaster manager needing a flood map right now, what's the problem with relying on this image alone?
Section 6 — Water-Sensitive Spectral Indices
Raw bands are hard to threshold reliably. Instead we compute normalized difference indices, which are more robust to lighting/atmospheric differences:
- NDWI = (Green − NIR) / (Green + NIR) — classic water index
- MNDWI = (Green − SWIR1) / (Green + SWIR1) — generally better for turbid/sediment-laden flood water
Compute MNDWI for the pre-flood composite (we'll need this one later, to identify permanent water).
mndwi_pre = s2_pre.normalizedDifference(['B3', 'B11']).rename('MNDWI')
Compute NDWI and MNDWI for the flood-period composite.
ndwi_post = s2_post.normalizedDifference(['B3', 'B8']).rename('NDWI')
mndwi_post = s2_post.normalizedDifference(['B3', 'B11']).rename('MNDWI')
Visualize the flood-period NDWI. Brighter blue = more water-like.
vis_ndwi = {'min': -0.5, 'max': 0.5, 'palette': cm.palettes.ndwi}
map_ndwi = geemap.Map(center=[9.25, 76.55], zoom=10)
map_ndwi.addLayer(ndwi_post, vis_ndwi, 'NDWI (flood period)')
map_ndwi.addLayerControl()
map_ndwi
- Compare this to the RGB flood image from Section 5. Does NDWI make the flooded area easier to see than the raw RGB composite did?
- Where does cloud cover still break the index?
Section 7 — SAR (Sentinel-1): Pre-Flood Image
Sentinel-1 is a radar satellite — it sends out its own microwave pulses and measures what bounces back, in two polarizations (VV, VH). Because it supplies its own energy source and uses wavelengths that pass through clouds and rain, it can image the surface any time, regardless of weather. Smooth surfaces like open water reflect radar away from the sensor (like a mirror), so water shows up dark (low backscatter) in SAR.
Load the Sentinel-1 GRD collection, filtered to the interferometric-wide (IW) mode with both VV and VH polarizations.
s1_collection = (ee.ImageCollection('COPERNICUS/S1_GRD')
.filterBounds(aoi)
.filter(ee.Filter.eq('instrumentMode', 'IW'))
.filter(ee.Filter.listContains('transmitterReceiverPolarisation', 'VV'))
.filter(ee.Filter.listContains('transmitterReceiverPolarisation', 'VH'))
.select(['VV', 'VH']))
Define a simple speckle filter. SAR images are inherently noisy (speckle) — a focal median smooths this out.
def speckle_filter(image, radius=30):
filtered = image.focal_median(radius, 'circle', 'meters')
return ee.Image(filtered.copyProperties(image, ['system:time_start']))
Build the pre-flood SAR composite: median of the pre-flood scenes, clipped to the district, then speckle-filtered.
s1_pre = speckle_filter(s1_collection.filterDate(PRE_START, PRE_END).median().clip(aoi))
Visualize the VH band. Dark = low backscatter = likely water.
vis_sar = {'bands': ['VH'], 'min': -25, 'max': 0}
map_s1_pre = geemap.Map(center=[9.25, 76.55], zoom=10)
map_s1_pre.addLayer(s1_pre, vis_sar, 'Sentinel-1 VH Pre-flood')
map_s1_pre.addLayerControl()
map_s1_pre
- Compare the dark areas here to the water bodies you saw in the optical pre-flood image (Section 4). Do rivers, backwaters and the lake line up?
Section 8 — SAR (Sentinel-1): Flood-Period Image
Build the flood-period SAR composite the same way.
s1_post = speckle_filter(s1_collection.filterDate(POST_START, POST_END).median().clip(aoi))
Visualize it.
map_s1_post = geemap.Map(center=[9.25, 76.55], zoom=10)
map_s1_post.addLayer(s1_post, vis_sar, 'Sentinel-1 VH Flood period')
map_s1_post.addLayerControl()
map_s1_post
- Unlike the optical flood image, is this one fully covered with no gaps?
- Can you spot new dark areas (new water) that weren't dark in the pre-flood SAR image?
Section 9 — Optical vs. SAR, Side by Side, Same Flood Dates
This is the key "appreciate the capability" moment: the exact same date range, the same district, two different sensors.
compare_map = geemap.Map(center=[9.25, 76.55], zoom=10)
left_layer = geemap.ee_tile_layer(s2_post, vis_rgb, 'Optical RGB (flood)')
right_layer = geemap.ee_tile_layer(s1_post, vis_sar, 'SAR VH (flood)')
compare_map.split_map(left_layer=left_layer, right_layer=right_layer)
compare_map
- Drag the slider across the map. Where optical is cloud-free, do the two sensors agree on where the water is?
- Where optical is cloudy, SAR is your only source of information. For rapid disaster response, which would you trust first?
Section 10 — SAR Change Detection: Where Did Backscatter Drop?
We can make the flood signal even more explicit by directly subtracting the pre-flood VH image from the flood-period VH image. A big negative change means backscatter dropped sharply — a strong sign of new standing water.
vh_change = s1_post.select('VH').subtract(s1_pre.select('VH')).rename('VH_change')
vis_change = {'bands': ['VH_change'], 'min': -8, 'max': 8, 'palette': ['blue', 'white', 'red']}
map_change = geemap.Map(center=[9.25, 76.55], zoom=10)
map_change.addLayer(vh_change, vis_change, 'VH change (flood - pre)')
map_change.addLayerControl()
map_change
- Blue areas = backscatter dropped (likely new flooding). Red areas = backscatter increased — why might that happen? (Think about flooded vegetation causing double-bounce scattering, a well-known SAR quirk.)
Section 11 — Collecting Training Samples for Random Forest
Random Forest is a supervised classifier — it needs labeled examples of water (class 1) and non-water (class 0) to learn from. Rather than guessing coordinates by hand (error-prone, and it doesn't generalize if you change the study area), we'll let the pre-flood MNDWI tell us where permanent water already is, then draw many random sample points from inside and outside that mask, spread across the whole district. This is a standard, defensible way to generate training labels automatically.
First, define what counts as "water" in the pre-flood image: any pixel with MNDWI above 0.
permanent_water_mask = mndwi_pre.gt(0).rename('class')
Now draw 80 random points from each class (water / non-water), using Earth Engine's stratifiedSample. geometries=True keeps the actual point coordinates so we can map them.
training_fc = permanent_water_mask.stratifiedSample(
numPoints=80,
classBand='class',
region=aoi,
scale=10,
seed=42,
geometries=True
)
How many points did we get, and how are they split between the two classes?
print('Total training points:', training_fc.size().getInfo())
print('Class histogram:', training_fc.aggregate_histogram('class').getInfo())
Output
Total training points: 160
Class histogram: {'0': 80, '1': 80}
- Roughly balanced classes are good for Random Forest. If one class had 10x more points than the other, what problem might that cause?
Let's actually SEE where these training points are, color-coded by class, on top of the flood-period SAR image.
water_pts = training_fc.filter(ee.Filter.eq('class', 1))
nonwater_pts = training_fc.filter(ee.Filter.eq('class', 0))
training_map = geemap.Map(center=[9.25, 76.55], zoom=10)
training_map.addLayer(s1_post, vis_sar, 'Reference: SAR VH (flood)')
training_map.addLayer(water_pts.style(color='00BFFF', pointSize=4), {}, 'Training: water (class 1)')
training_map.addLayer(nonwater_pts.style(color='FF4500', pointSize=4), {}, 'Training: non-water (class 0)')
training_map.addLayerControl()
training_map
- Do the blue (water) points land on rivers/backwaters, and the orange (non-water) points on dry land?
- These labels come from the PRE-FLOOD image, so they represent permanent water, not the flood itself. Keep that in mind — we'll come back to it in Section 16's accuracy discussion.
Optional: Add Your Own Training Points Interactively
If you want to practice drawing your own training data instead of (or in addition to) the automatic sample, use the draw-point tool on the map below, then run the tagging cell under it.
collect_map = geemap.Map(center=[9.25, 76.55], zoom=10)
collect_map.add_basemap('Esri.WorldImagery')
collect_map.addLayer(s1_post, vis_sar, 'Reference: SAR VH (flood)')
collect_map.addLayerControl()
collect_map.add_draw_control()
collect_map
Set up two buttons — "Label as Water" and "Label as Non-Water" — that tag whatever point you most recently drew on the map above:
my_points = [] # holds extra ee.Feature points you tag below
water_button = widgets.Button(description="Label as Water (Class 1)", button_style='info')
non_water_button = widgets.Button(description="Label as Non-Water (Class 0)", button_style='warning')
output = widgets.Output()
def on_button_clicked(button, class_value):
with output:
output.clear_output()
if not collect_map.draw_features:
print('Please draw a point on the map first using the drawing tool (point icon), then click a label button.')
return
last_drawn_feature = collect_map.draw_features[-1]
if last_drawn_feature.geometry().type().getInfo() != 'Point':
print('Only point features can be labeled. Please draw a point.')
return
ee_point = last_drawn_feature.set('class', class_value)
my_points.append(ee_point)
print(f'Tagged a point as class={class_value}. Total custom points: {len(my_points)}')
collect_map.draw_features.clear()
water_button.on_click(lambda b: on_button_clicked(b, 1))
non_water_button.on_click(lambda b: on_button_clicked(b, 0))
display(widgets.VBox([water_button, non_water_button, output]))
print("Instructions: Draw a point on the map above using the drawing tools (point icon), then click 'Label as Water' or 'Label as Non-Water'. Repeat for more points.")
- If you added your own points, we'll merge them into the automatic sample in the next section.
Section 12 — Extracting Pixel Values at the Training Locations
So far training_fc only has locations and class labels — no actual band/index values yet. Now we extract the real pixel data at each point from the flood-period imagery, which is what the Random Forest will actually learn from.
If you added any points interactively in Section 11, merge them in now (safe to run even if you added none).
if len(my_points) > 0:
training_fc = training_fc.merge(ee.FeatureCollection(my_points))
print(f'Merged in {len(my_points)} of your own points. New total: {training_fc.size().getInfo()}')
else:
print('No extra points added — using the automatic sample only.')
Stack every band/index we might want into one multi-band image: optical bands, the two indices, and the two SAR polarizations.
optical_bands = ['B3', 'B4', 'B8', 'B11']
index_bands = ['NDWI', 'MNDWI']
sar_bands = ['VV', 'VH']
combined_bands = optical_bands + index_bands + sar_bands
stacked_post = (s2_post.select(optical_bands)
.addBands(ndwi_post)
.addBands(mndwi_post)
.addBands(s1_post.select(sar_bands)))
Now sample the stacked image at every training point location.
samples = stacked_post.select(combined_bands).sampleRegions(
collection=training_fc,
properties=['class'],
scale=10,
tileScale=4
)
print('Samples with valid pixel values:', samples.size().getInfo())
Output
Samples with valid pixel values: 150
- The sample count may be slightly lower than the number of training points — some points fall under cloud-masked pixels in the flood-period optical image and get dropped. This is itself a reminder of the cloud problem from Section 5.
Let's actually look at a few rows of real feature values, so the abstraction of "sampling" becomes concrete numbers.
preview_df = geemap.ee_to_df(samples.limit(8))
preview_df
Output
| B11 | B3 | B4 | B8 | MNDWI | NDWI | VH | VV | class |
|---|---|---|---|---|---|---|---|---|
| 0.1819 | 0.0938 | 0.0708 | 0.3147 | -0.3196 | -0.5408 | -12.25 | -7.11 | 0 |
| 0.1870 | 0.0863 | 0.0569 | 0.3459 | -0.3685 | -0.6006 | -12.79 | -5.67 | 0 |
| 0.2167 | 0.1155 | 0.0807 | 0.3582 | -0.3046 | -0.5124 | -14.75 | -6.33 | 0 |
| 0.3277 | 0.2347 | 0.2299 | 0.3225 | -0.1654 | -0.1576 | -14.18 | -6.82 | 0 |
(First 4 of 8 preview rows shown — all class 0 / non-water in this slice of the sample.)
- Scan the VH and MNDWI columns against the class column. Do class=1 (water) rows tend to have lower VH (more negative) and higher MNDWI than class=0 rows?
- This is literally the pattern the Random Forest is about to learn — just automated and extended across EVERY combination of bands, not just one or two.
Section 13 — Splitting into Training and Test Sets
If we train the classifier on ALL our samples and then check its accuracy on those same samples, we're grading our own homework — the accuracy will look artificially good. Instead, we hold back a portion of the samples the model never sees during training, and only use those for evaluation.
Add a random number to each sample, then split 70% into training and 30% into testing.
samples_split = samples.randomColumn('random_value', seed=1)
train_samples = samples_split.filter(ee.Filter.lt('random_value', 0.7))
test_samples = samples_split.filter(ee.Filter.gte('random_value', 0.7))
print('Training samples:', train_samples.size().getInfo())
print('Test samples:', test_samples.size().getInfo())
Output
Training samples: 109
Test samples: 41
- Why 70/30, and not something more extreme like 99/1 or 50/50?
- What's the trade-off between giving the model more training data vs. having a trustworthy test set?
Section 14 — Random Forest Model 1: Optical Bands Only
We'll train three separate Random Forest models so we can directly compare optical-only, SAR-only and combined performance. Starting with optical.
optical_model_bands = optical_bands + index_bands # B3, B4, B8, B11, NDWI, MNDWI
Train the classifier using only the training split.
classifier_optical = ee.Classifier.smileRandomForest(numberOfTrees=50).train(
features=train_samples,
classProperty='class',
inputProperties=optical_model_bands
)
Apply it to the full flood-period image to produce a wall-to-wall water/non-water map.
classified_optical = stacked_post.select(optical_model_bands).classify(classifier_optical)
Visualize the result.
vis_water = {'min': 0, 'max': 1, 'palette': ['white', '#1f78ff']}
map_optical_result = geemap.Map(center=[9.25, 76.55], zoom=10)
map_optical_result.addLayer(classified_optical, vis_water, 'Classified water — Optical model')
map_optical_result.addLayerControl()
map_optical_result
- Does the water extent look plausible given what you saw in the optical imagery earlier?
- Remember this model had no information in cloud-covered areas during training — where might it be guessing?
Section 15 — Random Forest Model 2: SAR Bands Only
classifier_sar = ee.Classifier.smileRandomForest(numberOfTrees=50).train(
features=train_samples,
classProperty='class',
inputProperties=sar_bands
)
classified_sar = stacked_post.select(sar_bands).classify(classifier_sar)
map_sar_result = geemap.Map(center=[9.25, 76.55], zoom=10)
map_sar_result.addLayer(classified_sar, vis_water, 'Classified water — SAR model')
map_sar_result.addLayerControl()
map_sar_result
- Compare this to the optical-only result. Any differences near cloud-affected zones from Section 5?
Section 16 — Random Forest Model 3: Combined Optical + SAR
classifier_combined = ee.Classifier.smileRandomForest(numberOfTrees=50).train(
features=train_samples,
classProperty='class',
inputProperties=combined_bands
)
classified_combined = stacked_post.select(combined_bands).classify(classifier_combined)
map_combined_result = geemap.Map(center=[9.25, 76.55], zoom=10)
map_combined_result.addLayer(classified_combined, vis_water, 'Classified water — Combined model')
map_combined_result.addLayerControl()
map_combined_result
- Combining data sources doesn't automatically guarantee the best result — it depends on whether the extra bands add real information or just noise. Keep this question in mind for the next section, where we measure it properly.
Section 17 — Accuracy Assessment
Now we evaluate each model on the test set it never saw during training — this is the number that actually means something.
Classify the test samples themselves (not the whole image) with each model, so we can compare their predictions to the true labels.
tested_optical = test_samples.classify(classifier_optical)
tested_sar = test_samples.classify(classifier_sar)
tested_combined = test_samples.classify(classifier_combined)
Build a confusion matrix for each model and pull out accuracy + kappa.
results = []
for label, tested in [('Optical-only', tested_optical), ('SAR-only', tested_sar), ('Combined', tested_combined)]:
error_matrix = tested.errorMatrix('class', 'classification')
results.append({
'Model': label,
'Test accuracy': round(error_matrix.accuracy().getInfo(), 3),
'Kappa': round(error_matrix.kappa().getInfo(), 3),
})
accuracy_df = pd.DataFrame(results)
accuracy_df
Output
| Model | Test Accuracy | Kappa |
|---|---|---|
| Optical-only | 0.878 | 0.759 |
| SAR-only | 0.902 | 0.802 |
| Combined | 0.927 | 0.853 |
- Which model generalizes best to unseen pixels?
- If SAR-only underperforms, remember our training LABELS came from an optical index (MNDWI) — could that bias the comparison in optical's favor? How would you design a fairer test?
Plot the comparison as a bar chart:
fig, ax = plt.subplots(figsize=(6, 4))
ax.bar(accuracy_df['Model'], accuracy_df['Test accuracy'], color=['#4C72B0', '#DD8452', '#55A868'])
ax.set_ylabel('Test accuracy (held-out samples)')
ax.set_ylim(0, 1)
ax.set_title('Random Forest test accuracy by data source')
for i, v in enumerate(accuracy_df['Test accuracy']):
ax.text(i, v + 0.02, f'{v:.2f}', ha='center')
plt.tight_layout()
plt.show()
Section 18 — Isolating New Flood Water from Permanent Water
Our classifier labels all water the same way — rivers, the lake, and newly flooded fields all come back as class 1. For a flood map, we usually want to know specifically what's newly underwater. We can isolate that by subtracting the permanent water mask (from Section 11, based on the pre-flood image) from the flood-period classification.
new_flood_combined = classified_combined.eq(1).And(permanent_water_mask.eq(0)).rename('new_flood')
vis_new_flood = {'min': 0, 'max': 1, 'palette': ['white', '#d7191c']}
map_new_flood = geemap.Map(center=[9.25, 76.55], zoom=10)
map_new_flood.addLayer(permanent_water_mask.selfMask(), {'palette': ['#1f78ff']}, 'Permanent water (pre-flood)')
map_new_flood.addLayer(new_flood_combined.selfMask(), vis_new_flood, 'NEW flood water')
map_new_flood.addLayerControl()
map_new_flood
- This red layer is arguably the single most useful output of the whole workshop: it answers "what area is flooded that wasn't flooded before?" — the question an emergency manager actually needs answered.
Section 19 — How Many km² Flooded?
Sum the pixel area of the NEW flood mask across the whole district.
new_flood_area_image = new_flood_combined.multiply(ee.Image.pixelArea())
new_flood_stats = new_flood_area_image.reduceRegion(reducer=ee.Reducer.sum(), geometry=aoi, scale=10, maxPixels=1e9)
new_flood_km2 = new_flood_stats.getNumber('new_flood').divide(1e6).getInfo()
print(f'New flood extent: {new_flood_km2:,.1f} km2')
Output
New flood extent: 99.9 km2
For context, also compute the total water extent (permanent + new) during the flood.
total_water_image = classified_combined.eq(1).multiply(ee.Image.pixelArea())
total_water_stats = total_water_image.reduceRegion(reducer=ee.Reducer.sum(), geometry=aoi, scale=10, maxPixels=1e9)
total_water_km2 = total_water_stats.getNumber('classification').divide(1e6).getInfo()
print(f'Total water during flood (permanent + new): {total_water_km2:,.1f} km2')
print(f'Share of the district newly underwater: {100 * new_flood_km2 / area_km2:.1f}%')
Output (illustrative — total_water_km2 will print from your own run)
Total water during flood (permanent + new): <value from your run> km2
Share of the district newly underwater: 7.6%
- Does the percentage of the district newly flooded match what you'd expect from a "worst flood in a century" event?
- What are the main sources of error in this estimate (classifier mistakes, cloud gaps, speckle, the 10-day-wide date windows)?
Section 20 — Wrap-Up Discussion
- We used a Random Forest with a hand-picked 50 trees and default settings. What would you try next to improve accuracy (more/different bands, more trees, more training points, a different classifier)?
- Our training labels came entirely from thresholding the pre-flood optical image. What's circular or risky about that, and how would you validate it independently (hint: field surveys, high-resolution imagery, official flood extent reports)?
- We used fixed 7–23 day composite windows. How would the analysis change for a flood that develops and recedes within 2–3 days?
- If you had to repeat this entire workflow for a different river basin or country tomorrow, which parts of this notebook would you have to change, and which parts would work unmodified?
Exporting Your Results
To save the combined flood map to your Google Drive as a GeoTIFF:
export_task = ee.batch.Export.image.toDrive(
image=new_flood_combined.toByte(),
description='Kerala_2018_Alappuzha_NewFloodExtent',
folder='flood_mapping_workshop',
region=aoi,
scale=10,
maxPixels=1e9
)
# Uncomment to start the export (runs in the background on Earth Engine's servers):
# export_task.start()
# print('Export started — check the Tasks tab at code.earthengine.google.com/tasks')
Summary
- Optical (Sentinel-2) is intuitive and detailed but unreliable during floods due to cloud cover.
- SAR (Sentinel-1) sees through clouds and rain, making it the most reliable source for rapid disaster response, at the cost of being noisier and less intuitive to interpret visually.
- Training labels were generated automatically via stratified sampling on a pre-flood water index, avoiding error-prone manual coordinate guessing.
- Random Forest, trained on those labels, combined optical and SAR features to reach 92.7% test accuracy on held-out samples — beating either data source alone (Optical: 87.8%, SAR: 90.2%).
- Subtracting a pre-flood permanent-water mask from the flood-period classification isolates the specific, actionable answer: how much NEW area flooded (99.9 km², ~7.6% of the district).
- The entire workflow runs for free in Google Colab via Google Earth Engine — no local data download, no GIS software required.