---
title: "Regional growth, convergence, and spatial spillovers in India: A reproducible view from outer space"
# REGION-compatible author format
author:
- name: "Carlos Mendez"
affiliations:
- name: "Nagoya University"
city: "Nagoya"
country: "Japan"
orcid: "0000-0001-7978-2815"
email: "carlosmendez777@gmail.com"
corresponding: true
- name: "Sujana Kabiraj"
affiliations:
- name: "Shiv Nadar University"
city: "Greater Noida"
country: "India"
- name: "Jiaqi Li"
affiliations:
- name: "Nagoya University"
city: "Nagoya"
country: "Japan"
abstract: |
Using satellite nighttime light data as a proxy for economic activity, Chanda and Kabiraj (2020, World Development) studied regional growth and convergence across 520 districts in India.
Adopting a reproducible open-science approach, this article builds on their work by extending their main findings on three empirical fronts.
First, we illustrate regional convergence patterns using an interactive tool for satellite imagery visualization.
Second, we assess the degree of spatial dependence in their main econometric specification.
Third, we employ a spatial Durbin model to measure the role of spatial spillovers in the convergence process.
Our results indicate that spatial spillovers increase the estimated speed of regional convergence.
In our fully specified model, spillovers raise the convergence speed from about 3.0% to about 5.2% per year, which shortens the half-life of regional disparities from roughly 23 to 13 years.
We close by illustrating how the same reproducible workflow extends beyond economic output, with an exploratory analysis of how luminosity relates to cultural participation across Indian states.
keywords:
- Regional convergence
- Spatial dependence
- Spatial Durbin model
- Nighttime lights
- Reproducible research
- India
bibliography: references.bib
# REGION journal metadata
received: "February 5, 2026"
accepted: "September 19, 2026"
jauthor: "C. Mendez, S. Kabiraj, J. Li"
# JEL classification codes (optional but recommended for economics journals)
jel:
- R11 # Regional Economic Activity: Growth, Development
- R12 # Size and Spatial Distributions of Regional Economic Activity
- C21 # Cross-Sectional Models; Spatial Models
# Year of the copyright notice printed by the REGION template
jyear: 2026
# Journal-assigned at acceptance. As of 2026-09-20 REGION has not issued a
# volume, issue or OJS number, so jvol/jnum/ojsnum stay commented and
# partials/title.tex falls back to its defaults: "Volume 1, Number 1" on the
# title page and in the running foot, and the placeholder DOI
# 10.18335/region.v??i??.???. That is intentional -- the journal overwrites
# them at typesetting. jpages is the measured page count of this final build,
# and _quarto.yml resets the page counter to 1 to match it.
# jvol: 11
# jnum: 1
jpages: "1--26"
# ojsnum: 456 # Assigned by journal for DOI generation
---
```{=latex}
\keywords{Regional convergence, Spatial dependence, Spatial Durbin model, Nighttime lights, Reproducible research, India}
\jel{R11, R12, C21}
```
## Introduction
Regional growth and convergence matter most in federal states such as India.
In such settings, spatial inequalities can threaten social cohesion and political stability.
Consistent economic data, however, are rarely available below the state level.
Satellite nighttime lights have relaxed that constraint by serving as a proxy for economic activity.
Using those data, @chanda_kabiraj_district_convergence document convergence across 520 Indian districts between 1996 and 2010.
However, their analysis does not account for spatial spillovers.
Such spillovers can arise from the diffusion of technology and from the movement of resources across neighboring districts.
This article extends that analysis along three dimensions.
The first is replication: we adopt the data, the sample, and the conditioning variables of @chanda_kabiraj_district_convergence without modification, so that any difference in the estimates reflects the spatial modeling alone.
The second is substantive: we build an interactive visualization of regional convergence, test formally for spatial dependence, and estimate a spatial Durbin model.
The model shows that catch-up across Indian districts has a significant neighborhood dimension.
The spatial spillover effect of initial luminosity is negative and statistically significant.
Districts whose neighbors started from a lower base therefore grew faster than their own initial conditions alone would predict.
The third contribution is methodological, and it is the primary one.
This article shows that a comprehensive spatial econometric study can be implemented using only open-source software and cloud computing.
Our workflow covers every step, from visualizing the satellite imagery and constructing the neighborhood structure to testing for spatial dependence, modeling the spillovers, and running the robustness checks.
That workflow delivers the main empirical result of the article.
Accounting for spatial spillovers raises the estimated speed of regional convergence from about 3.0% to about 5.2% per year.
The implied half-life of regional disparities therefore falls from roughly 23 years to 13.
The same workflow extends beyond economic output, and we illustrate that reach with an exploratory analysis of luminosity and cultural participation across Indian states.
Our results suggest that luminosity is positively associated with media-based cultural consumption and negatively associated with community-based participation.
This pattern is consistent with luminosity tracking electrification and urbanization rather than economic output alone.
Every figure and table in this article is produced by one of the six cloud-based computational scripts and notebooks listed in @tbl-notebooks.
The five analytical notebooks are written in Python and run in Google Colaboratory, so a reader needs no local installation and no licensed software: clicking *Run all* downloads the data and re-executes the analysis.
The remaining notebook, N1, hosts the interactive visualization and its Google Earth Engine source code.
The notebooks, the data, and the manuscript source are in the project repository at <https://github.com/quarcs-lab/project2025s-py>, and their rendered versions are embedded in the online edition at <https://quarcs-lab.github.io/project2025s-py/>.
| Notebook | Content |
|:---------|:--------|
| N1. View from outer space | Interactive luminosity map with Earth Engine script|
| N2. Regional convergence | Luminosity growth on initial luminosity, with speed and half-life |
| N3. Spatial dependence | Weight matrices, Global Moran's I, and LISA cluster maps |
| N4. Spillover modeling | OLS and spatial Durbin estimates, impacts, and convergence speeds |
| N5. Robustness | Model 4 under six alternative weight matrices |
| N6. Spatial culture | Luminosity and cultural participation across Indian states |
: Computational notebooks and scripts underlying this article {#tbl-notebooks}
The rest of this article is organized as follows.
Section 2 describes the data and the methods, and Section 3 presents the empirical results.
Section 4 discusses the policy implications, the extension to multiple periods, the analysis of cultural participation, and new directions for research.
Section 5 offers concluding remarks.
## Data and methods
### Data: Nighttime lights as a proxy for economic activity
Satellite images of nighttime light provide a proxy for economic activity below the national level.
@henderson_storeygard_weil_lights pioneered this approach in economics.
Their work exploits the empirical correlation between the intensity of artificial light observed from space and economic output on the ground.
@chen_nordhaus_luminosity_gdp show that the proxy is most informative where official statistics are of low quality.
Nighttime light (hereafter NTL) data have therefore become a standard tool for the study of growth and convergence across national and subnational regions.
NTL data have been applied to a wide range of questions in development economics.
One line of work studies the effect of decentralization on regional convergence [@adhikari_dhital_decentralization].
A second line adjudicates between national accounts and household surveys, and it finds that national accounts track the light data more closely [@pinkovskiy_salaimartin_lights_poverty].
A third line estimates GDP per capita and spatial inequality at the global scale [@lessmann_seidel_regional_inequality].
The Indian literature is equally varied, and it covers welfare programs, colonial institutions, demonetization, and the pandemic.[^india]
@chakravarty_dehejia_gst document large regional disparities in India and caution that the goods and services tax may widen them further.
Taken together, these studies show that NTL data are well suited to the analysis of subnational disparities.
[^india]: Nighttime lights have been used in India to analyze the impact of public welfare programs [@cook_shah_nregs], the effects of colonial institutions [@jha_talathi_colonial_india], the impact of demonetization [@chanda_cook_demonetization], and the effects of COVID-19 [@beyer_jain_sinha_covid].
Our empirical design follows @chanda_kabiraj_district_convergence.
The dependent variable is the growth rate of nighttime lights per capita.
The primary explanatory variable is the initial level of nighttime lights per capita.
We construct both variables for 520 districts in India from data released by the National Geophysical Data Center (NGDC).
These data record observations from the DMSP-OLS satellites over the period from 1996 to 2010.
To mitigate top-coding in brightly lit areas, the NGDC also released a radiance-calibrated series for eight specific years, which combines high magnification settings for low-light regions with low magnification settings for bright ones.
Six of those eight years fall inside our 1996--2010 sample window, and they are the dates our dataset holds.
In this article, we use the radiance-calibrated series throughout.[^1]
[^1]: Following @chanda_kabiraj_district_convergence, our sample consists of 520 districts (out of a possible 593).
There were 593 districts in 2001, which rose to 640 districts in the 2011 census.
To match districts across the two census files, we merged newly split districts back with their parent districts in the 2011 census.
Eight of the 47 new districts were created by splitting areas from multiple parent districts.
Those new districts along with their multiple-origin districts were dropped from our sample.
We dropped all the districts in the state of Assam, where more than half of the new districts were created in that manner.
The 73 districts not in our sample are those of Assam and Tripura, the union-territory districts outside Delhi, the eight by which the nine districts of Delhi are collapsed to one, and 28 multiple-origin districts and their parents.
### A reproducible workflow and interactive visualization
We implement the entire study in a reproducible workflow.
Jupyter notebooks combine executable code, explanatory narrative, and computational output in a single document [@kluyver_jupyter], which is the practical expression of literate programming [@knuth_literate].
We document the data processing, the estimation, and the visualization in notebooks that run on a single Python kernel.
Each result can therefore be traced to the code that produced it.
Quarto then extends those notebooks into a manuscript framework [@allaire_quarto].
Every figure and table is embedded from an identified notebook cell, and one source file generates the HTML, PDF, Word, and JATS XML versions without discrepancies between them.
The toolchain is open source, and the data and code are public [@peng_reproducible].
Each notebook also runs in the cloud on Google Colaboratory, which requires no local installation.
Readers can therefore re-execute the whole pipeline in a browser and without proprietary software.
Satellite data have opened new dimensions of economic measurement [@donaldson_storeygard_remotesensing].
Interactive visualization complements that measurement, because it reveals spatial and temporal heterogeneity that static maps obscure.
Google Earth Engine lowers the computational barriers to such visualization [@gorelick_gee].
Its browser-based environment requires no local software, and its catalog already holds the DMSP-OLS and VIIRS nighttime lights collections [@tamiminia_gee_review].
Readers can therefore track light intensity across regions over time, and they can see where initially dim regions brightened fastest.
That pattern is the visual counterpart of the convergence relationship estimated in Section 3.
### Regional convergence modeling
Neoclassical growth theory predicts that poor regions grow faster than rich ones.
Diminishing returns to capital accumulation make the per capita growth rate of a region negatively correlated with its initial endowment [@solow_growth].
Regions that share technology and preferences should therefore converge to a common steady state in the long run.
We test that prediction within the growth regression framework of @barro_sala_convergence.
@eq-matrix-form-abs states the convergence process in its simplest unconditional form.
$$
\boldsymbol{g_t} = \beta_1 \boldsymbol{x_{t-1}} + \boldsymbol{\varepsilon_t}
$$ {#eq-matrix-form-abs}
where $\boldsymbol{g_t}$ is an $N\text{-by-}1$ vector of per-capita NTL growth for the $N$ regions over the period $t$.
The vector $\boldsymbol{x_{t-1}}$ collects the initial (log) per-capita NTL of the same regions, and $\boldsymbol{\varepsilon_t}$ collects idiosyncratic error terms.
The parameter $\beta_1$ measures the direction and the strength of regional convergence.
A negative value implies that regions with lower initial NTL levels grow faster, which is the convergence hypothesis.
The estimated convergence coefficient translates into two quantities that are easier to interpret [@barro_sala_convergence].
The first is the annual speed of convergence, and the second is the half-life.
Because the dependent variable is average annual growth over $T$ years, the implied convergence speed is $\lambda = -\ln(1 + \beta_1 T)/T$.
The half-life is $\ln(2)/\lambda$, which is the time required to close half of the initial gap to the steady state.
In our application $T = 14$ years.
We apply the same transformation to the total effect of initial luminosity in the spatial models.
The unconditional framework contains no control variables, so it imposes a common steady state on every region.
That assumption is restrictive, because regions differ in geography, in socioeconomic conditions, and in policy implementation.
We therefore add state fixed effects, which absorb state-specific institutions and policies.
We also add geo-climatic controls and district-specific conditions related to demographics, human capital, and infrastructure.
@eq-matrix-form-cond summarizes the resulting conditional convergence framework.
$$
\boldsymbol{g_t} = \beta_1 \boldsymbol{x_{t-1}} + \boldsymbol{X_t} \boldsymbol{\alpha} + \boldsymbol{\varepsilon_t}
$$ {#eq-matrix-form-cond}
Here, the matrix $\boldsymbol{X_t}$ is an $N\text{-by-}k$ collection of observations on control variables for each district in our sample, and it includes the state fixed effects.
The vector $\boldsymbol{\alpha}$, with dimensions $k\text{-by-}1$, collects the associated regression coefficients.
The conditioning set is adopted unchanged from @chanda_kabiraj_district_convergence.
That choice is deliberate, because the present analysis extends that study rather than replacing it.
Holding the controls fixed ensures that any difference between our estimates and their estimates comes from the spatial specification and not from a different choice of covariates.
The 16 district-level controls fall into three groups.[^controls]
Geography and agro-climate capture slow-moving determinants of the steady state that do not vary over the sample period.
Demographics and human capital capture the factor endowments through which districts accumulate.
Infrastructure captures access to the electricity grid and to markets.
This last group matters here, because luminosity is itself partly a measure of electrification.
The conditional models therefore compare districts that began the period with similar levels of grid access.
The demographic, human-capital, and electrification controls are measured in 1996, at the start of the growth period.
They are therefore predetermined with respect to subsequent growth.
The paved-roads variable is measured in 2000.
The full list of control variables, with their measurement years and summary statistics, is given in Table A.1 of @chanda_kabiraj_district_convergence.
[^controls]: Geography and agro-climate comprise agricultural suitability, rainfall, temperature, a malaria ecology index, terrain ruggedness, latitude, and distance to the coast.
Demographics and human capital comprise the rural population share, log population density, the scheduled caste and scheduled tribe shares, the working population share, and the literate and higher-educated shares.
Infrastructure comprises the share of households with an electricity connection and the log number of households with access to paved roads.
We estimate both specifications on a single long difference.
The dependent variable is average annual growth between 1996 and 2010, and the regressors are measured at the start of that period.
This is the design of @chanda_kabiraj_district_convergence, and we retain it.
The design identifies the average rate at which districts converged over fourteen years.
It is silent, however, on how that rate evolved within the period.
We return to this limitation, and to what a panel design would require, in the discussion.
### Spatial dependence testing
Spatial dependence is the tendency for observations at nearby locations to be more similar than observations at distant locations.
It can arise from spatial spillovers, from shared geographic conditions, or from inter-regional economic linkages [@anselin_spatial_analysis].
When it is present and left unmodeled, standard regressions yield biased parameter estimates and invalid inference.
We therefore test for spatial dependence with the Global Moran's I statistic and with Local Indicators of Spatial Association (LISA).
We apply both tests to the main variables of @eq-matrix-form-cond.
The Global Moran's I statistic quantifies the overall degree of spatial clustering among geographic units [@moran_spatial_autocorrelation].
It compares the deviation of each observation from the mean with the deviations of its neighbors.
@eq-moran gives the formula.
$$
I=\frac{n}{\sum_i \sum_j w_{ij}} \cdot \frac{\sum_i \sum_j w_{ij} \, z_i \, z_j}{\sum_i z_i^2}
$$ {#eq-moran}
where $z_i$ is the deviation of observation $i$ from the mean, and $w_{ij}$ is the spatial weight that connects units $i$ and $j$.
The term $n$ denotes the total number of observations.
The statistic typically ranges from $-1$ to $+1$.
Positive values indicate that similar values tend to be located near each other.[^moranrange]
We assess significance with a permutation-based inference approach.[^perm]
[^moranrange]: Values near zero suggest spatial randomness, and negative values indicate spatial dispersion, where dissimilar values are neighbors.
[^perm]: The observed statistic is compared against a reference distribution generated by randomly reassigning values across locations.
The global statistic does not reveal where the clusters and outliers are located.
@anselin_lisa therefore proposed Local Indicators of Spatial Association, which decompose the global statistic into a contribution from each observation.
@eq-lisa defines the local statistic for observation $i$.
$$
I_i = z_i \sum_j w_{ij} \, z_j
$$ {#eq-lisa}
where the summation is over the neighbors of $i$, as defined by the spatial weight matrix.
Each observation is then classified into one of four categories, according to how its value relates to the values of its neighbors.
The High-High and Low-Low categories identify spatial clusters, and the High-Low and Low-High categories identify spatial outliers.[^lisa]
We assess significance with conditional permutation tests.
The LISA cluster maps report only the significant locations, typically at $p < 0.05$.
[^lisa]: High-High (HH) indicates a high-value location surrounded by high-value neighbors, and Low-Low (LL) indicates a low-value location surrounded by low-value neighbors.
High-Low (HL) is a spatial outlier where a high-value location is surrounded by low-value neighbors, and Low-High (LH) is the opposite spatial outlier.
The spatial weight matrix $\mathbf{W}$ determines which districts are treated as neighbors, so its specification is central to the analysis.
We employ a six-nearest-neighbor (6NN) matrix.
It sets $w_{ij} = 1$ if district $j$ is among the six nearest districts to $i$ by centroid distance, and $w_{ij} = 0$ otherwise.[^centroids]
Each district therefore has the same number of neighbors, even though district sizes vary widely across the country.
We then row-normalize the matrix, as is standard practice, so that each row sums to one.
The spatial lag $\mathbf{W} \mathbf{x}$ is then the average value among the neighbors of a district.
The choice of $k = 6$ balances two concerns.
It is large enough to capture local spatial interactions, and it is small enough to exclude distant districts that are unlikely to exert economic influence.
The data also support this choice.
Among the seven weight matrices considered in @tbl-altw, the 6NN specification attains the lowest AIC, which indicates the best fit to the spatial structure of the data.
We nevertheless treat the definition of the neighborhood as an open question.
In Section 3 we re-estimate the preferred model under six alternative weight matrices, which range from contiguity-based schemes to inverse-distance schemes.
Those results, reported in @fig-altw and @tbl-altw, show that the direct and total convergence effects do not depend on this choice.
[^centroids]: Centroid distances are measured in the geographic coordinates of the source file rather than in a projected metric, and rebuilding the matrix from a metric projection leaves 96.4% of the neighbor links unchanged.
Evidence of significant spatial dependence would justify the use of spatial econometric methods.
@ertur_koch_spatial_growth argue that the spatial Durbin model (SDM) accounts for such dependence in the convergence process.
@fischer_spatial_mrw derives the same specification as the reduced form of a spatially interdependent growth model.
The SDM also separates the direct effects of district characteristics from the indirect effects that operate through spatial channels [@lesage_pace_spatial_econometrics].
### Spatial spillover modeling
Our spatial spillover modeling builds on the spatially augmented Solow model of @ertur_koch_spatial_growth, which was developed for countries.
It also builds on the regional Mankiw-Romer-Weil extension of that model by @fischer_spatial_mrw.
Both models account for technological interdependence across economies.
We apply this framework to an economy of $N$ subnational regions.
Each region is characterized by the Cobb-Douglas production function with constant returns to scale in @eq-production.
$$
Y_{i t}=A_{i t} K_{i t}^{\alpha_{K}} H_{i t}^{\alpha_{H}} L_{i t}^{1-\alpha_{K} -\alpha_{H}}
$$ {#eq-production}
where $Y_{it}$ is output, $K_{it}$ is physical capital, $H_{it}$ is human capital, $L_{it}$ is the labor force, and $A_{it}$ is the level of technological knowledge for region $i$ at time $t$.
The parameters $\alpha_K$ and $\alpha_H$ are the output elasticities with respect to physical and human capital.
Technological knowledge is not exogenous in this framework, because it depends on internal and on external factors.
@eq-technology states that relation.
$$
A_{i t}=\Omega_{t} k_{i t}^{\theta} h_{i t}^{\phi} \prod_{j \neq i}^{N} A_{j t}^{\rho W_{i j}}
$$ {#eq-technology}
The specification captures three components of technological progress:
- An exogenous component ($\Omega_t$) that represents the common stock of knowledge across regions.
- An embodied component ($k_{i t}^{\theta} h_{i t}^{\phi}$) that reflects technology embedded in physical and human capital per worker.
- A spatial component ($\prod_{j \neq i}^{N} A_{j t}^{\rho W_{i j}}$) that captures technological interdependence between regions.
The third component is absent from the standard Solow model, and it is the source of the spatial spillovers in growth.
From this framework @ertur_koch_spatial_growth derived a spatial Durbin model that accounts for spillovers in the convergence process, with the human-capital terms following @fischer_spatial_mrw.
The model augments the conditional specification of @eq-matrix-form-cond with spatial lags of the dependent variable, of initial luminosity, and of the control variables, and can be written compactly as @eq-matrix-form-sdm:
$$
\boldsymbol{g_t} = \beta_1 \boldsymbol{x_{t-1}} + \boldsymbol{X_t} \boldsymbol{\alpha} + \beta_2 \boldsymbol{W} \boldsymbol{x_{t-1}} + \boldsymbol{W} \boldsymbol{X_t} \boldsymbol{\gamma} + \rho \boldsymbol{W} \boldsymbol{g_t} + \boldsymbol{\varepsilon_t}
$$ {#eq-matrix-form-sdm}
The vectors $\boldsymbol{g_t}$, $\boldsymbol{x_{t-1}}$ and $\boldsymbol{\varepsilon_t}$, the matrix $\boldsymbol{X_t}$, and the parameters $\beta_1$ and $\boldsymbol{\alpha}$ are defined as in @eq-matrix-form-cond above.
The new terms are the spatial lags: $\boldsymbol{W} \boldsymbol{x_{t-1}}$ and $\boldsymbol{W} \boldsymbol{g_t}$ are $N\text{-by-}1$ vectors of neighborhood averages of initial luminosity and of growth, while $\boldsymbol{W} \boldsymbol{X_t}$ is the corresponding $N\text{-by-}k$ matrix for the control variables.
The coefficients $\beta_2$, $\boldsymbol{\gamma}$ and $\rho$ measure how strongly each of these neighborhood averages enters the growth of a district.[^rho]
[^rho]: We denote the spatial autoregressive parameter by $\rho$ to avoid any confusion with $\lambda$, which denotes the implied speed of convergence throughout this article.
## Results
### Regional convergence: An interactive exploration from outer space
Before presenting the regression results, we illustrate growth and convergence in nighttime lights visually, through an interactive map of India, a convergence scatterplot, and case studies of some of the poorest regions of the country.[^2]
[^2]: The interactive web application is available at [https://bit.ly/india-rc-ntl](https://bit.ly/india-rc-ntl).
It was developed using Google Earth Engine, and the source code is available in the [View from outer space](https://quarcs-lab.github.io/project2025s-py/notebooks/c01_view_from_space.html) notebook.
. `<br>`{=html}`\protect\newline`{=latex} Source: Authors' visualization using pre-processed luminosity images from the Earth Observation Group (NOAA/NCEI). See [View from outer space](https://quarcs-lab.github.io/project2025s-py/notebooks/c01_view_from_space.html) notebook for source code.](images/luminosity_map.png){#fig-map}
@fig-map presents static maps of luminosity in the initial and final years, 1996 and 2010, captured from our interactive application.
The maps show a noticeable increase in brightness across most of India by 2010, which, since nighttime lights proxy for economic activity, reflects economic growth over the period.
@fig-convergence then relates per-capita growth in nighttime lights to initial per-capita luminosity, and the relationship is inverse: the estimated $\beta$-convergence coefficient of about $-0.02$ implies an annual speed of convergence of roughly 2.3% and a half-life of about 30 years, consistent with the seminal finding of @barro_sala_convergence.
Because $\beta$-convergence need not imply a narrowing of the distribution, we also check $\sigma$-convergence: the cross-district standard deviation of log per-capita luminosity falls from 1.09 in 1996 to 0.95 in 2010, a reduction of 13.5%, so spatial disparities did narrow over the period.
{{< embed notebooks/c02_regional_convergence_sc.ipynb#fig-convergence >}}
@fig-map2 displays the change in luminosity over the study period in three economically disadvantaged states.
Bihar is the poorest of them, with a per-capita income at 31.2% of the national average, while Uttar Pradesh and Chhattisgarh stand at 55.3% and 59.9% respectively.[^3]
All three remain among the poorest states, but all three brightened noticeably over the period.
[^3]: The figures are relative per capita income for 2000--01, the census year closest to the start of our study period, measured as per capita net state domestic product divided by per capita net national income.
They are taken from Table 4 of @sanyal_arora_state_performance, Working Paper EAC-PM/WP/31/2024.
. `<br>`{=html}`\protect\newline`{=latex} Source: Authors' visualization using pre-processed luminosity images from the Earth Observation Group (NOAA/NCEI). See [View from outer space](https://quarcs-lab.github.io/project2025s-py/notebooks/c01_view_from_space.html) notebook for source code.](images/luminosity_map2.png){#fig-map2 width=75%}
### Spatial dependence is a feature of the convergence process
Before the formal econometric analysis, we examine the spatial distribution of the variables under study.
@fig-chorophleths maps initial luminosity in 1996 and the subsequent growth rate of luminosity over the 1996--2010 period across the 520 districts in our sample.
{{< embed notebooks/c03_spatial_dependence_lisa.ipynb#fig-chorophleths >}}
Panel (a) shows initial luminosity concentrated in specific geographic corridors, with higher values clustered in western and north-western India and markedly lower values across large parts of eastern and northern India.
Panel (b) shows a pattern that is largely the inverse: districts that were initially bright grew more slowly, and districts that were initially dim grew faster.
This spatial inversion provides a first visual indication that regional convergence in India may have an important spatial dimension.
A formal analysis of spatial dependence requires the neighborhood of each district, which the 6NN weight matrix described above supplies.
@fig-wmatrix6nn visualizes that connectivity as a network, in which each node represents a district centroid and each edge connects a district to one of its six nearest neighbors.
{{< embed notebooks/c03_spatial_dependence_lisa.ipynb#fig-Wmatrix6nn >}}
Spatial dependence can be understood through two complementary concepts: the overall degree of clustering in the data, and the specific locations of statistically significant clusters and spatial outliers.
The Global Moran's I statistic captures the first of these concepts.
In our data, the Moran's I for initial luminosity per capita is 0.73 ($p = 0.001$), which indicates strong positive spatial autocorrelation: districts with high (low) initial luminosity tend to be surrounded by districts that are also bright (dim).
The Moran's I for luminosity growth is 0.60 ($p = 0.001$), which confirms that the growth rates of neighboring districts are also significantly correlated.
We now turn to the LISA statistics, which locate the significant clusters and outliers behind these global measures.
@fig-dependence-initial presents the Moran scatterplot and LISA cluster map for initial luminosity per capita.
HH clusters mark the most luminous regions and their bright neighbors, while LL clusters highlight contiguous areas of low luminosity.
These are predominantly in eastern and northern India, in Bihar, eastern Uttar Pradesh, and Jharkhand, together with the Himalayan north and the northeast.
{{< embed notebooks/c03_spatial_dependence_lisa.ipynb#fig-dependence-initial >}}
The same analysis applied to luminosity growth rates reveals a pattern that is largely the inverse of the initial clusters, as @fig-dependence-growth shows.
No district that formed an HH cluster in initial luminosity forms one in luminosity growth: some appear instead as LL clusters, and most lose significance.
Part of the initially dimmest cluster, in turn, emerges as an HH cluster in growth.
That inversion is the spatial signature of the convergence process, and it motivates the use of spatial econometric methods.
{{< embed notebooks/c03_spatial_dependence_lisa.ipynb#fig-dependence-growth >}}
### Evidence of spatial spillovers in regional convergence
The regression results of @tbl-models provide evidence of both unconditional and conditional convergence across districts in India, under ordinary least squares (OLS) and under the spatial specifications alike.
The direct effects, which represent within-district convergence, remain negative and of similar order across specifications, ranging from -0.020 to -0.026.
This consistency across specifications and estimation methods suggests that poorer districts are catching up to their wealthier counterparts, even after controlling for district characteristics and state-level fixed effects.
```{=latex}
\begingroup\footnotesize\setlength{\tabcolsep}{3pt}
```
{{< embed notebooks/c04_spillover_modeling_6nn.ipynb#tbl-models >}}
```{=latex}
\endgroup
```
In the OLS specifications, the direct effect is -0.020 in Model 1 and -0.022 in Model 2, and it strengthens to -0.025 once district characteristics are added in Models 3 and 4.
Omitting these characteristics therefore attenuates the estimated speed of convergence, and the spatial Durbin direct effects match this pattern in the conditional models, at -0.026 in Model 3 and -0.025 in Model 4.
State fixed effects absorb unobserved state-level heterogeneity, such as differences in governance and institutional quality, which tightens the standard errors on the spatial impacts and moderates the spatial total effect rather than strengthening it further.
The spatial Durbin model also reveals spillover effects that OLS does not capture, and these indirect effects, which measure the influence of the initial conditions of neighboring districts on the growth rate of a district, follow a notable pattern across specifications.
In the unconditional models (Models 1 and 2), the estimated indirect effects are negative but small and statistically insignificant, at -0.001 in both.[^erratic]
Once district-level controls are introduced (Models 3 and 4), the indirect effects become both larger in magnitude and statistically significant at the 10% level, reaching -0.015 in Model 3 and -0.012 in our most comprehensive Model 4.
The spatial channels of convergence therefore become discernible only after accounting for district-specific characteristics.
[^erratic]: Without controls, the strong residual spatial dependence confounds the spillover estimates with omitted variables.
The spatial autoregressive parameter also reaches about 0.80 in Model 1, so its decomposition is the most sensitive of the four to how the spatial multiplier is computed.
The total impact of initial conditions on growth, combining both direct and spillover effects, is substantially larger when we account for spatial dependence.
In our fully specified model (Model 4), the total convergence effect in the spatial Durbin model (-0.037) is approximately 51% larger in magnitude than the OLS estimate (-0.025).
Translated into an annual speed of convergence, this raises the implied speed from about 3.0% under OLS to about 5.2% under the SDM, and it shortens the implied half-life from roughly 23 to 13 years.
The gap is even larger in Model 3 (SDM total -0.041 against OLS -0.025, or 67%), but the Model 3 and Model 4 spillover estimates are statistically indistinguishable, so we anchor our interpretation on the preferred Model 4.
These differences indicate that in this setting conventional non-spatial approaches understate the speed of regional convergence by failing to capture the additional convergence channels created through spatial spillovers.
{{< embed notebooks/c04_spillover_modeling_6nn.ipynb#tbl-speed >}}
@tbl-speed reports the implied speeds and half-lives for all four specifications.
The spatial premium is concentrated in the conditional models.
In Models 1 and 2 the OLS and SDM speeds are nearly identical, whereas in Models 3 and 4 the SDM raises the implied speed from 3.0% to 6.2% and 5.2%, shortening the half-life from 23 years to 11 and 13 years respectively.[^uncondspeed]
[^uncondspeed]: The unconditional speeds are 2.3% against 2.6% in Model 1 and 2.6% against 2.7% in Model 2.
The Akaike Information Criterion (AIC) consistently favors the spatial Durbin model over OLS across all four specifications, and the most comprehensive model (Model 4 SDM) achieves the lowest AIC value of -2501 compared to -2465 for the corresponding OLS.[^aic]
These results suggest that regional convergence in India operates not only through district-specific factors but also through spatial interactions between neighboring districts, so that ignoring those interactions yields both underestimated convergence speeds and inferior model performance.
[^aic]: The improvement from incorporating spatial structure is most pronounced in the unconditional specification, where the AIC drops by 347 units when moving from OLS to SDM in Model 1, reflecting the large amount of spatial dependence left unmodeled by ordinary regressions.
A word of caution is in order about how the indirect effects should be read.
The spatial Durbin model identifies spatial dependence in reduced form: it establishes that the growth of a district covaries with the initial conditions of its neighbors, conditional on its own characteristics, but it does not establish why.
Several distinct processes generate the same reduced-form pattern, since neighboring districts share geographic and agro-climatic conditions, are often governed by the same state and so inherit common institutional and policy environments that our state fixed effects absorb only in part, are connected by infrastructure whose effects do not stop at administrative boundaries, and are exposed to any omitted determinant of growth that is itself spatially clustered.
The measurement issues discussed below point in the same direction, since blooming spreads light across district boundaries and can raise the apparent correlation between neighbors for reasons unrelated to economic interaction.
We therefore read the indirect effects as evidence that spatial structure is a feature of the convergence process, and not as an estimate of the causal effect of the conditions of one district on the growth of another.
Establishing which channels are at work, and whether they are causal, would require research designs built for that purpose, a point we return to in the discussion.
### How these estimates compare with the convergence literature
The speeds reported above can be placed against three benchmarks.
The first is other work on Indian districts using the same proxy.
@beyer_yamada_district_banks report 2.4% absolute and 3.7% conditional convergence for India over 2013--2019, against our own 2.3% and 3.0%.
That two independent applications of the same measurement strategy recover similar non-spatial rates, despite different periods and different generations of the sensor, is reassuring about the proxy.
The second benchmark is convergence measured from luminosity elsewhere.
@basher_etal_bangladesh_lights find absolute convergence of 4.57% per year across the 64 districts of Bangladesh, roughly double the 2.3% we obtain on the same unconditional basis.
@carrington_jimenez_ecuador_lights report a speed near 2% for Ecuadorian provinces, and @kacou_africa_convergence finds multiple equilibria across 49 African countries.
Our own unconditional non-spatial estimate therefore sits near the familiar 2% regularity, while the conditional spatial estimates lie well above it.[^twopercent]
A meta-analysis of some 600 published estimates nevertheless warns against treating that regularity as a law [@abreu_etal_meta_convergence].
These comparisons locate our estimates rather than validate them.
[^twopercent]: The regularity is about 2% per year across 1,528 regions in 83 countries [@gennaioli_etal_growth_regions], and similar across regions of the United States, Japan, and five European countries [@salaimartin_regional_cohesion].
The meta-analytic objection is that estimates from different specifications are drawn from different populations, so there is no single natural rate to be estimated.
The third benchmark bears most directly on what this article adds.
That convergence estimates are sensitive to spatial dependence is established, but the direction of the adjustment is not.
Our own rise from 3.0% to 5.2% agrees with @diop_africa_spillovers, who concludes that non-spatial estimates may be considerably understated.
Others find the opposite: lower speeds for European regions [@isla_castillo_etal_eu_cohesion], neighbor effects that slow convergence across Indonesian districts [@santos_marquez_etal_indonesia_spatial], and divergence rather than convergence across Indian states [@lolayekar_mukhopadhyay_india_spatial].
The substantive point is therefore not that spatial dependence matters, which is already known, but that here it works in the direction of faster measured convergence, and that the scale of analysis is part of what determines the answer.[^signs]
[^signs]: In our own case the indirect effect carries the same negative sign as the direct effect, so the two reinforce one another and the total effect exceeds the OLS coefficient in magnitude.
Where the initial conditions of neighbors enter with the opposite sign, or where the spatial terms absorb variation the non-spatial model had attributed to initial luminosity, the same exercise reduces the estimated speed.
### Robustness to alternative spatial weight matrices
The spillover results above use the 6NN spatial weight matrix defined in Section 2.
To assess their sensitivity to this choice, we re-estimate the preferred Model 4 under six alternative specifications and compare them with the 6NN baseline.
The alternatives are four- and eight-nearest neighbors, queen and rook contiguity, and inverse distance and inverse-distance-squared, the last two applied within a distance band.
{{< embed notebooks/c07_alternative_w_matrices.ipynb#fig-altw >}}
```{=latex}
\begingroup\footnotesize
```
{{< embed notebooks/c07_alternative_w_matrices.ipynb#tbl-altw >}}
```{=latex}
\endgroup
```
The direct and total convergence effects are robust to the choice of spatial weights, and the spillover component is consistently negative but more sensitive (@fig-altw and @tbl-altw).
The direct (within-district) convergence effect is essentially unchanged across all seven specifications, ranging from -0.024 to -0.025 and always significant at the 1% level.
The total convergence effect likewise remains negative and significant throughout, ranging from -0.032 to -0.041, with the 6NN baseline (-0.037) near the middle of this range.
The indirect (spillover) component is negative in every specification and statistically significant under five of the seven, which is a consistent if weaker pattern.
This stability indicates that the direct and total convergence effects are not an artifact of the particular neighbor definition.
The results reported so far treat luminosity as a proxy for economic output, which is the use that the convergence literature has made of it.
The discussion that follows asks first what the magnitude of these estimates implies for regional policy, and what rests on the use of a single growth period.
It then turns to a second use of the same data, since nighttime lights also record electrification, urbanization, and broader socioeconomic conditions, and the reproducible workflow assembled here can be pointed at those dimensions as readily as at output.
That possibility is pursued in an exploratory way, asking how luminosity relates to cultural participation across Indian states.
## Discussion
### What the magnitude of the effects implies for regional policy
The difference between the two readings of Model 4 is large enough to matter for policy.
Under the OLS estimate, a district closes half of the gap to its steady state in about 23 years; under the spatial Durbin estimate, in about 13 years.
Over the fourteen years of our sample, that is the difference between closing roughly 34% and roughly 52% of the initial gap.
The first figure describes catch-up over a generation; the second, catch-up within the horizon over which development programs are designed, funded, and evaluated.
Under the faster rate, the returns to well-targeted support should appear within a normal evaluation window.
The results also speak to the level at which such policies should operate.
We find convergence across districts over 1996--2010, whereas @lolayekar_mukhopadhyay_india_spatial, working with Indian states over a longer period, find divergence.
Catch-up visible at the district level and absent at the state level is consistent with the district being the more informative unit of intervention, which the Aspirational Districts Programme of India already assumes.[^adp]
The spatial results add a qualification, since a lagging district surrounded by other lagging districts is in a different position from one adjacent to a growing cluster.
A targeting rule that ranks districts one at a time will not register that difference.
[^adp]: The program was launched in 2018 to direct additional resources to the least developed districts of the country.
The two studies also differ in outcome variable and period, so the contrast between them is suggestive rather than decisive.
Spillovers cut both ways, however, and the Indian evidence on place-based policy is a caution as much as an encouragement.
@hasan_etal_backward_districts find that a tax exemption for industrially backward districts raised firm entry and employment partly by displacing activity from neighboring districts that narrowly missed qualifying, with effects that did not outlast the program.
Displacement of that kind would appear in our framework as a spatial relationship between neighbors, which is not a spillover on which coordinated regional policy can be built.
What the satellite record does support is measurement.[^shenoy]
The readings above therefore hold only if the indirect effects reflect genuine interaction between districts, and the reduced-form evidence is consistent with that but does not establish it.
[^shenoy]: @shenoy_place_based_policy uses nighttime lights to evaluate a large place-based package in Uttarakhand and finds that luminosity rises sharply on the targeted side of the new state border, which shows that the outcome we study is responsive enough to register policy at this scale.
### The single growth period and what a panel would require
The estimates reported above rest on a single growth period, and the cost of that design should be stated plainly.
One long difference cannot trace how convergence evolved within the period, cannot show whether the spatial spillovers were stronger in some years than in others, and cannot separate long-run convergence from what was specific to these fourteen years, which in India included the maturing of the post-1991 reforms.
Examining those dynamics lies beyond an article that adds a spatial dimension to an existing cross-sectional result.
The immediate obstacle is in any case the data, since district population is observed only at the decennial censuses.[^denominator]
Interpolated denominators touch only the endpoints of a long difference, whereas in a panel they would enter every observation, and would do so in the time dimension that a panel exists to exploit.
[^denominator]: The dataset we adopt from @chanda_kabiraj_district_convergence holds radiance-calibrated luminosity at six dates between 1996 and 2010, so per-capita luminosity at the intermediate dates would rest on interpolated denominators.
What such a panel would be worth is the more interesting question, and here two problems compound.
The first predates nighttime lights: panel estimates of convergence run far above cross-sectional ones, and @caselli_etal_convergence_debate obtain a rate near 10% per year against the 2% that cross-sections deliver.[^panelgap]
How that gap should be read remains contested, so a panel would give not a better-identified version of our estimate but an estimate of something different.
The second problem is specific to the proxy, since luminosity discriminates differences across places far better than changes over time [@chen_nordhaus_viirs_crosssection; @asher_shrug_india], and over short intervals the ratio of signal to measurement error falls sharply.[^shorthorizon]
The periods compared should therefore be long, and harmonized DMSP--VIIRS series now make a sequence of non-overlapping long differences feasible across roughly three decades [@li_etal_harmonization], which is the most promising temporal extension of this work.
[^panelgap]: @barro_convergence_modernisation argues that country fixed effects generate a misleadingly high convergence rate, and that combining post-1960 with post-1870 evidence returns a figure close to the conventional 2%.
Meta-analytic evidence finds that corrections for unobserved heterogeneity raise estimated convergence systematically [@abreu_etal_meta_convergence].
[^shorthorizon]: A short-horizon panel would therefore pair a specification that is hard to interpret with a proxy operating where it is weakest.
@chakrabarty_mukherjee_covid_convergence estimate district convergence across the pandemic period, and how such estimates should be read depends on how much of the short-run movement in lights reflects output rather than measurement.
### An exploratory extension: Luminosity and cultural participation
Nighttime lights capture not just GDP but a composite of electrification, urbanization, infrastructure, and broader socioeconomic conditions [@henderson_storeygard_weil_lights; @mellander_etal_lights_gdp].
In this section, we examine the association between luminosity and cultural participation patterns across Indian states.
A growing literature argues that regional cultural attitudes constitute an independent factor in economic development, not merely a by-product of income or institutions.
In particular, @tubadji_cultural_entropy applies the Culture-Based Development framework, whose central claim is a directional one.
Cultural attitudes and the participation patterns that reveal them are treated as a factor shaping local development, rather than as an outcome that development produces.
That direction also marks the limit of what the exercise below can establish.[^direction]
[^direction]: A cross-sectional association is consistent with either direction, and with both operating at once.
What it can do is establish that the two are related, and describe the form that relationship takes.
For India, the National Sample Survey (NSS) 47th Round (July--December 1991) surveys household cultural participation, and nighttime lights for the adjacent year 1992 are available for 32 states and union territories.
Following the Culture-Based Development framework, we reduce these survey items to six dimensions and aggregate them to state means.[^nss]
This part of the analysis uses the Consistent and Corrected Nighttime Lights series rather than the radiance-calibrated composites used above, because it is the product that covers 1992.
Given the small sample ($N$ = 32) and the leverage that small territories at the extremes of the distribution exert on linear correlation measures, we report Spearman rank correlations throughout.
@fig-culture-scatter presents the relationship between log nighttime lights per capita and the two cultural dimensions that show statistically significant associations.
[^nss]: The six dimensions are live cultural performance, cultural telecast (TV/media), socio-cultural participation, cultural heritage and religion, live cultural shows, and sports.
They are extracted from the survey items by principal-components analysis.
See the [Spatial culture](https://quarcs-lab.github.io/project2025s-py/notebooks/c06_spatial_culture.html) notebook for details and extended analyses.
{{< embed notebooks/c06_spatial_culture.ipynb#fig-culture-scatter >}}
Cultural telecast (TV/media) is positively associated with luminosity, with a Spearman rank correlation of 0.370 ($p = 0.037$), so more luminous states tend to rank higher on that dimension.
Socio-cultural participation, in contrast, is negatively associated, with a Spearman rank correlation of $-0.404$ ($p = 0.022$), so less luminous states tend to rank higher on it.
We refer to these two dimensions as media-based cultural consumption and community-based participation, respectively.
This suggests that less urbanized regions rely more on collective, in-person forms of cultural participation than on media infrastructure.
@fig-culture-lisa maps the spatial distribution of the two cultural variables through the lens of a LISA analysis.
The Low-Low cluster of media-based cultural consumption in the northeast coincides with the Low-Low cluster in the nighttime lights data, whereas the High-High clusters do not coincide: media-based cultural consumption forms one in the southern states, and community-based participation forms one in the northeast.
The remaining four cultural dimensions show no statistically significant association with luminosity, and these findings rest on a small cross-section of 32 states observed at a single point in time, so larger datasets covering more regions and periods would be needed to establish whether the patterns are robust.
{{< embed notebooks/c06_spatial_culture.ipynb#fig-culture-lisa >}}
Two connections to the main analysis are worth drawing out, both offered as interpretation rather than as findings.
The first concerns what luminosity measures.
A positive association with media-based cultural consumption and a negative one with community-based participation are what one would expect if luminosity tracks electrification and urbanization rather than economic output alone.
The cultural results are, in that sense, a small piece of evidence about the composite nature of the proxy on which the convergence analysis rests, and they bear on the measurement issues discussed below.
The second connection concerns spatial structure.
Cultural participation is not distributed randomly across states, since both dimensions examined here are significantly spatially autocorrelated.
Factors of this kind, spatially clustered and correlated with development, are precisely what a spatial Durbin model absorbs into its indirect effects when they are not measured directly.
We cannot test this with state-level data for a single adjacent period, and we do not claim that cultural participation drives the spillovers estimated in the results section.
What the exercise does is make concrete what a reduced-form spatial model may be capturing, and it suggests that the Culture-Based Development framework offers one route to testing the question directly, should comparable data become available at the district level and over time.
### New research directions
Three kinds of gap limit what this literature can currently establish: the data themselves, the geography they cover, and the theories they have been used to test.
The first is well documented, since the radiance-calibrated DMSP-OLS series we adopt from @chanda_kabiraj_district_convergence, and on which this literature rests [@henderson_storeygard_weil_lights; @chen_nordhaus_luminosity_gdp], is top-coded in bright urban cores, uncalibrated on board, and blurred across boundaries by blooming [@gibson_etal_ntl_measurement; @abrahams_etal_deblurring].
Top-coding leaves the net bias in our convergence coefficient ambiguous in sign, while blooming raises the apparent correlation between neighbors and therefore bears directly on the spillover estimates.[^bias]
The Visible Infrared Imaging Radiometer Suite (VIIRS) and harmonized DMSP--VIIRS series address all three problems [@elvidge_etal_viirs; @li_etal_harmonization; @chiovelli_global_south], and multi-source methods improve measurement where lights alone are weak.[^multisource]
Re-estimating our specification on those data is the most immediate extension of the analysis reported here.
[^bias]: Top-coding compresses variation at the bright end of the distribution, so classical measurement error in initial luminosity attenuates the convergence coefficient toward zero, while the same error entering the growth rate with the opposite sign works in the direction of overstating convergence.
The net effect is ambiguous in sign, and we do not claim to sign it.
Measurement error is also unlikely to be homogeneous across the sample, because the association between luminosity and economic outcomes is stronger in urban than in rural areas and Indian districts differ considerably in how urbanized they are.
That the direct and total effects hold across all seven definitions of the neighborhood, and the indirect effect under five of the seven (@tbl-altw), indicates that they are not an artifact of any particular neighbor definition, but it does not rule out blooming, for which deblurring methods offer one route [@abrahams_etal_deblurring].
[^multisource]: VIIRS, operational since 2012, improves spatial resolution (roughly 750 meters against 2.7 kilometers), adds on-board radiometric calibration, and widens dynamic range, which avoids saturation in bright areas while detecting dim lights in rural settlements [@gibson_etal_ntl_measurement].
Lights have also been combined with daytime imagery for poverty prediction [@jean_etal_poverty_prediction], with land cover in agricultural areas [@keola_etal_lights_poverty], with multiple remote sensing indicators for subnational output [@chen_etal_turkey_viirs; @hussein_etal_vietnam_ml], and with survey data for multidimensional poverty [@theara_etal_cambodia_poverty].
The second gap concerns geography, and it is not that some regions have gone unstudied.
A single instrument observes the whole planet under a common protocol, which makes measurements comparable in a way that national accounts are not [@henderson_storeygard_weil_lights], but what it records is light rather than economic activity.
Roughly one fifth of the settlement footprint of the planet emits no detectable radiance, and unlit settlements are concentrated in Africa [@mccallum_unlit_settlements].
Africa is nonetheless well covered by this literature [@kacou_africa_convergence], as are China [@glawe_mendez_china_luminosity], Indonesia [@miranti_mendez_indonesia], Thailand [@tipayalai_mendez_thailand], and Turkey [@ursavas_mendez_turkey].
The real gap is that the places where lights carry least information, rural and sparsely electrified areas, are the places where administrative statistics are weakest and a satellite proxy would be most valuable.[^africa]
Closing it calls less for further applications of the same proxy than for validation against ground data in those settings.
[^africa]: Africa accounts for 39% of unlit settlements worldwide, rising to 65% when only rural settlement areas are considered [@mccallum_unlit_settlements].
Across 34 sub-Saharan African countries the association between harmonized luminosity and household wealth is pronounced in urban areas and considerably weaker in rural ones [@jhamb_africa_nightlights], and systematic assessments of which product to use, and where, reach similar conclusions [@gibson_etal_ntl_measurement].
The third gap concerns theory, since luminosity is now applied well beyond the question asked here.
Within convergence analysis, our framework yields an average effect rather than heterogeneity across the distribution, and it captures spillovers in reduced form, gaps that distribution dynamics and quasi-experimental designs such as that of @asher_novosad_rural_roads are built to fill.[^clubs]
Beyond convergence, lights now serve the study of city size [@ch_martin_vargas_city_size], electrification and the provision of public goods [@min_etal_electrification], armed conflict [@witmer_lights_conflict], the cultural participation theorized by @tubadji_cultural_entropy, and the local effect of policies and shocks at fine spatial scales [@bluhm_mccord_small_geographies].
One caveat governs all of these uses: lights discriminate differences across places far better than changes over time, so evaluations over short horizons require validation against local ground data.[^timeseries]
The design used here, a fourteen-year horizon, rests mostly on cross-sectional variation, which is what lights currently support best.
[^clubs]: The relevant tools are distribution dynamics [@quah_twin_peaks] and the regression-tree methods developed for identifying multiple growth regimes [@durlauf_johnson_convergence_clubs], which can reveal whether Indian districts converge to a single steady state or form distinct clubs.
Combining them with spatial econometrics could expose spatial poverty traps or growth poles that average estimates render invisible [@rey_spatial_dynamics].
[^timeseries]: VIIRS predicts cross-sectional GDP better than time-series GDP [@chen_nordhaus_viirs_crosssection], and in India the elasticity of lights with respect to local outcomes is far lower in the time series than in the cross-section [@asher_shrug_india].
The relationship between lights and growth is also unstable across regions in both Brazil and India [@bickenbach_regional_gdp], and responsiveness to output varies with baseline income, density, and sector.
An evaluation may therefore fail to detect a real effect where lights react weakly, or attribute to a policy what is in fact heterogeneity in the relationship between light and output.
Taken together, these three gaps describe a research agenda rather than a list of caveats, and none of the work they imply requires starting from scratch.
The pipeline behind this article is public, documented, and executable in a browser, so a reader can open the relevant notebook, substitute data from another country or a later period, and re-run it.
That is the contribution we would most like this article to make: not an estimate of how quickly Indian districts converged between 1996 and 2010 alone, but a working template that lowers the cost of asking the same question elsewhere.
### Research reproducibility and open science
As @donaldson_storeygard_remotesensing note, methodological choices in processing remotely sensed data, such as sensor inter-calibration, can affect the empirical conclusions.
The same applies further down our own pipeline, at the construction of the spatial weight matrix.
Our own robustness analysis is a case in point: the estimated convergence effects survive seven definitions of the neighborhood, but that is something a reader can only confirm because the code that produced each of them is available.
Open-source tools and cloud platforms lower the barriers to that kind of verification, and cloud-based notebooks for processing nighttime lights remove the need for specialized local software [@mendez_patnaik_notebook].
Open data, version-controlled code, and cloud execution together make the full pipeline, from satellite imagery to econometric results, verifiable and extensible by the broader research community.
## Concluding remarks
This article re-examines regional convergence across Indian districts using satellite nighttime lights, interactive visualization, and spatial econometric modeling.
Our tests show that spatial dependence characterizes both the satellite data and the convergence process itself.
In the fully specified spatial Durbin model, the total convergence effect is about 51% larger than the conventional non-spatial estimate, and the implied half-life of regional disparities falls from roughly 23 years to 13.
Non-spatial models therefore understate the speed of convergence in this setting, although comparison with other regional literatures shows that the direction of the adjustment is specific to the setting.
The policy reading depends on what the indirect effects represent.
Genuine spillovers would mean that place-based interventions reach beyond the districts they target.
Displaced activity would mean that the same estimates overstate what those interventions add in aggregate.
Every step of the analysis is documented in Python notebooks, which run in the cloud and are embedded in the manuscript through the Quarto publishing framework.[^formats]
The code and the data are hosted in a public GitHub repository, and we follow these open-science practices in the hope that other applied studies will adopt them.
That same workflow allowed us to extend the analysis beyond economic output.
Relating luminosity to cultural participation across Indian states required one additional notebook.
Luminosity correlates positively with media-based cultural consumption and negatively with community-based participation.
Those exploratory findings are reproducible on the same terms as the convergence estimates.
The most valuable next steps are concrete: re-estimate this specification on the higher-resolution VIIRS data, validate luminosity against ground measurements in rural and sparsely electrified areas, and carry the same spatial framework to questions beyond convergence.
The convergence estimates reported here will eventually be superseded, and the workflow that produced them is built to be reused.
[^formats]: A single manuscript source generates multiple output formats.
The HTML version allows readers to engage with interactive visualizations, inspect the code of the computational notebooks, and re-evaluate the results in light of their source code.
## Acknowledgments {.appendix .unnumbered}
This research project was supported by JSPS KAKENHI Grant Number 24K04884.
During the preparation of this manuscript, the authors acknowledge the use of Claude Code (Anthropic) to assist with manuscript editing, computational notebook development, and research infrastructure setup.
After using this AI tool, the authors reviewed the outputs, confirmed their accuracy, and take full responsibility for the content of this publication.
## Conflict of interest {.appendix .unnumbered}
The authors declare no conflict of interest.
## Data and code availability {.appendix .unnumbered}
All data and computational code used in this study are available in the project repository at <https://github.com/quarcs-lab/project2025s-py>.
The interactive HTML version of this manuscript (<https://quarcs-lab.github.io/project2025s-py/>) embeds the computational notebooks, allowing readers to inspect the complete analytical pipeline from raw data to published results.
Every analytical notebook can be executed in the cloud using Google Colaboratory without any local software installation: readers open the notebook and run it, and the empirical results, figures, and tables of this article are regenerated from the raw data.
The computational notebooks are implemented entirely in Python.
An earlier implementation of the same analysis used R and Stata, and the present Python results reproduce it.
The AIC values differ by a few units because Stata and Python count parameters differently: Stata treats the error variance of the spatial model as a free parameter, and its information criterion for the fixed-effects regressions uses the rank of the robust covariance matrix, which two single-district states leave short by two.
The interactive Earth Engine application linked in the text is a convenience layer for visualizing the underlying imagery; its source code is included in the first notebook, and none of the results reported in this article depend on it.