Where development data is scarcest, it is needed most

The problem

  • A mayor planning a clinic has no current poverty number for their own territory.
  • A census is a snapshot, not a series. Bolivia has 339 municipalities, many too small or too remote to measure often.

The places that most need watching are the ones surveys reach last.

Lights, landscape, and people: three ways to see Bolivia

Three maps of Bolivia side by side on an identical frame. Left, nighttime lights: brightest at La Paz, Cochabamba and Santa Cruz, with the road network between them and scattered smaller settlements across the lowlands and highlands clearly visible; the light is tightly concentrated and thins out across the Chaco and altiplano. Centre, daytime embeddings as false colour: every pixel carries information, and the altiplano, Andean flank, Amazon lowlands and Chaco separate into distinct colour regions. Right, gridded population: inhabited cells form a thin scatter through the highland valleys and around the cities, leaving most of the territory empty.

All three for 2017, on the same frame · VIIRS nighttime lights · AlphaEarth embeddings, first three principal components as red, green and blue · GHS-POP gridded population

Development is visible from space — but which dimensions?

What is new

  • Google’s AlphaEarth turns a year of images into 64 numbers for each patch of land.
  • It learned on its own. Nobody showed it any social or economic data.
  • The lights add eight bands of brightness for the same place.

What we ask

  • We take 15 indices built from 62 indicators, for all 339 municipalities.
  • First: which goals can satellites read? Then: which can they map?
  • A random forest, tuned by Bayesian optimization, learns the link.
  • Every number in this talk is out of sample. No place helps predict itself.

Can satellites read development?

Part one · Prediction

Nighttime lights offer an opportunity to measure local development

Global map of Earth at night showing dense networks of artificial light across North America, Europe, India, eastern China, Japan, and major coastal and urban corridors, while sparsely populated regions remain largely dark.

NASA Black Marble · Suomi NPP VIIRS · global context, not the study raster

Daytime embeddings: a new way to measure local development

Conceptual workflow in which six daytime satellite-image tiles showing agriculture, a city, forest, wetlands, mountains, and dry terrain flow through the AlphaEarth Model into a sixty-four-dimensional satellite embedding represented by an eight-by-eight colored grid and a shortened numeric vector.

Conceptual illustration · multiple daytime observations become one compact embedding

A satellite sees pixels; an index describes a municipality

Flow chart in four columns. Left, three inputs: nighttime lights from the 2017 VIIRS annual composite with eight brightness bands, daytime embeddings from the 2017 AlphaEarth mosaic with 64 numbers per 10-metre pixel, and, in a dashed outline, the GHS-POP population surface. Both sensors feed one step, averaging the pixels inside each of the 339 municipal boundaries. That step forks into two: every pixel counted equally, the simple mean, and every pixel weighted by the people on it, with the population surface feeding the weighted branch by a dashed arrow. Both branches produce the output table, one row per municipality, 339 rows and 8 plus 64 predictors, built once under each scheme.

Rounded = a data product · sharp = a processing step · dashed = the population weights

What a goal is made of decides if satellites see it

The idea, before the evidence

  • Some goals are made of things you can see: housing, roads, land cover.
  • Others are made of things you cannot see: health, gender, justice.
  • So what a goal is made of should decide how well satellites read it.

On the raw average, the lights win the goals we care about most

Dumbbell plot of out-of-sample R-squared for all fifteen goal indices under a simple municipal average, comparing nighttime lights and daytime embeddings. Nighttime lights lead on the headline goals of poverty and energy; daytime embeddings lead on only eight of the fifteen.

Simple average · out of sample · 339 municipalities · daytime embeddings lead on only 8 of 15 — and trail on poverty and energy

Most of a municipality is empty — weight where people live

The methodological turn

Weight each pixel by the people on it and the embeddings rise from 0.29 to 0.41. The lights do not gain — they slip from 0.25 to 0.22.

Three maps of Bolivia side by side: nighttime lights, a false-colour map of daytime satellite embeddings, and gridded population. People and lights concentrate in a few valleys and cities while most of the territory is dark, empty land.

Weight where people live, and the embeddings pull ahead

Dumbbell plot of out-of-sample R-squared for all fifteen goal indices, comparing nighttime lights and daytime embeddings under population weighting. Daytime embeddings lead on fourteen of fifteen goals; both views approach zero for Peace, Justice and Strong Institutions.

Weighted by population · out of sample · daytime embeddings now lead on 14 of 15 goals (8 → 14) — though the lead clears fold noise on only 5 of the 15

The contribution is not one number — it is a map of legibility

  • Written into the physical fabric, so read well: poverty 0.69, energy 0.63, climate 0.63, hunger 0.58.
  • Assembled from social and administrative records, so barely read: institutions −0.05, gender 0.27, health 0.28.
  • Reduced Inequalities, at 0.57, is the best-read of the socially defined goals.
  • Life on Land, at 0.10, is the goal weighting costs us: its signal lives in the unsettled forest that weighting discards.

Legibility has a spatial signature

Slopegraph ranking the fifteen goal indices twice. On the left they are ranked by the global Moran's I of the actual index, from Climate Action at 0.65 down to Institutions at 0.09. On the right they are ranked by out-of-sample R-squared from the population-weighted daytime embeddings, from Poverty at 0.69 down to Institutions at minus 0.05. Lines connect each goal across the two rankings; most run roughly level and Spearman's rho is 0.67. Poverty, Energy and Climate Action are highlighted and occupy the top three places on the legibility side.

The goals that cluster most in space are the goals the daytime embeddings read best · fifteen goals, ranked twice · ranked, not to scale

Which goals can we map — and how well?

Part two · Monitoring

Together, the two views read poverty best of all

0.72 Poverty · the upper bound of a 0.54–0.72 range

We join the two views: 64 numbers from daytime images and 8 from nighttime lights, both weighted by population. Together they beat either one alone on poverty.

Three goals pass 0.60, our bar for drawing a map:

  • Poverty (SDG 1) 0.72
  • Energy (SDG 7) 0.64
  • Climate (SDG 13) 0.64

How you test decides what you get: 0.54 to 0.72

Dumbbell plot of the estimation range for the combined predictor across fifteen goals. Hollow markers show the department-transfer lower bound, filled markers the random five-fold upper bound. No Poverty spans 0.54 to 0.72; three goals clear the 0.60 threshold at the upper bound.

Each goal spans the two ways of testing · The three goals above the 0.60 line are the ones we map next

Under the hard test, only the fusion never collapses

Grouped horizontal bar chart of out-of-sample R-squared under department transfer for three goals and three feature sets. For No Poverty the combined predictor reaches 0.54 against 0.45 for nighttime lights and 0.44 for daytime embeddings. For Clean Energy it reaches 0.48 against 0.46 and 0.41. For Climate Action nighttime lights collapse to minus 0.28, daytime embeddings reach 0.46 and the combined predictor 0.44.

Tested in a department the model never saw — the bound that is not flattered by leakage

The combined view recovers the geography of poverty

Four cluster maps of Bolivia from local indicators of spatial association, comparing actual poverty clusters with those recovered from nighttime lights, daytime embeddings, and the combined predictor. Nighttime lights agree with 77% of all cluster classes and 53% of the actual hotspots and coldspots; daytime embeddings reach 78% and 80%, and the combined predictor 80% and 75%. The combined map reproduces the actual pattern most closely overall; global Moran's I is 0.50 actual, 0.33 lights, 0.60 embeddings, 0.50 combined.

Red: clusters of low poverty · Blue: poverty traps · Each panel reports its agreement with the actual map and its clustering strength

Clean energy: the lights read the level, not the pattern

Four cluster maps of Bolivia from local indicators of spatial association for the clean energy index, comparing actual clusters with those recovered from nighttime lights, daytime embeddings, and the combined predictor. Nighttime lights agree with 77% of all cluster classes but only 49% of the actual hotspots and coldspots; daytime embeddings reach 83% and 75%, and the combined predictor 82% and 70%. Global Moran's I is 0.43 actual, 0.28 lights, 0.49 embeddings, 0.40 combined.

Red: clusters of good access · Blue: traps of poor access · The lights reach r = 0.73 on the level, and half that on the pattern

Climate action: it finds the clusters — and overstates them

Four cluster maps of Bolivia from local indicators of spatial association for the climate action index. Nighttime lights agree with 64% of all cluster classes and 50% of the actual hotspots and coldspots; daytime embeddings reach 79% and 91%, and the combined predictor 81% and 92%. Global Moran's I is 0.65 actual, 0.50 lights, 0.81 embeddings and 0.84 combined, so both embedding-based views show markedly stronger spatial clustering than the actual map.

Red: clusters of strong action · Blue: traps of weak action · Best cluster recovery of the three goals — but both embedding views make the map look more clustered than it is

The lead is real, but smaller than it looks

Devil’s advocate

  • “14 of 15 is a point estimate.” — True: it clears noise for only five goals.
  • “It depends on the test.” — In a new region the flip is 7 → 8, not 8 → 14.
  • “Nearby places leak in.” — Hence both bounds. Poverty does worse than the local mean in four departments of nine.
  • “Climate action is circular.” — Partly: its deforestation component is measured from space. Poverty is not.
  • “Combining always wins?” — No: the embeddings alone find more hotspots on poverty and energy.

Satellites see the material geography of development

Let imagery track the goals it reads between survey rounds, and spend scarce survey effort on the goals it cannot read at all.

Together the two views read poverty at 0.54 to 0.72 out of sample and find its clusters and traps well enough to target scarce effort — but justice stays invisible, and health and gender barely register.

Validated on the 2017 cross-section. Carrying it forward between rounds is a projection, not a measurement.

Thank you

carlos-mendez.org