Finding Earth 2.0
Read this before citing a number

What this analysis cannot establish

A serious project states its limitations as prominently as its results. These are not disclaimers added after the fact — several were discovered while building the pipeline and changed how it works.

The Earth Similarity Index cannot separate Earth from Venus

Venus scores 0.92 on the Earth Similarity Index computed by this pipeline. Its high Bond albedo makes its equilibrium temperature cooler than Earth’s, and equilibrium temperature — not surface temperature — is the only temperature exoplanet catalogues provide. This is not a defect of this implementation; it is a property of the observations available for real exoplanets. No ESI computed from catalogue data can currently distinguish a temperate rocky world from a runaway-greenhouse one. Venus is carried through the entire pipeline as a labelled control specifically so this is visible in the results.

Methodological citation

Schulze-Makuch et al. (2011, Astrobiology 11, 1041) defines the Earth Similarity Index this project computes.

47% of catalogue masses were never measured

They are predictions from the radius via a mass–radius relation. Density and escape velocity computed from such a mass re-encode the radius rather than adding independent information, which would make the ESI appear to combine four independent properties while actually being driven by one. This project classifies mass evidence type explicitly and discounts inferred masses in the observational-confidence score.

Mass evidence, in full

Measured: 2,240. M sin i: 890. Inferred from radius: 2,975. Upper limit only: 218.

TRAPPIST-1 sits below the habitable-zone model’s validity floor

The Kopparapu et al. (2013) fit is stated valid for 2600–7200 K host effective temperatures. TRAPPIST-1’s host is 2566 K — 34 K below the floor. A strict reading excludes the most-studied temperate terrestrial system known from every habitable-zone count. This project reports results both ways: strictly (with TRAPPIST-1 planets carrying “undetermined” HZ status) and with an explicitly flagged extrapolation. Neither reading is hidden.

Methodological citation

Kopparapu et al. (2013, ApJ 765, 131; 2013 erratum, ApJ 770, 82) states the model’s 2600–7200 K validity range explicitly.

An Earth twin’s atmospheric signal is about 1 ppm

For a real nitrogen-oxygen atmosphere (mean molecular weight ≈29) around a Sun-like star, the transmission-spectroscopy amplitude is roughly an order of magnitude smaller than for a hydrogen-dominated atmosphere of the same scale height, and well below demonstrated JWST precision. Finding an Earth analogue and characterising its atmosphere are separated by a generation of instruments; nothing in this project promises otherwise.

The instrument gap

5,948 genuine transmission measurements exist across 104 planets — none of them Earth-sized in the habitable zone.

Discovery-method bias

The catalogue’s method distribution reflects instrument sensitivity, not the true underlying planet population. Transit surveys favour short periods and large radius ratios; radial-velocity surveys favour massive, close-in planets around bright quiet stars; direct imaging favours young, wide-separation giants. A temperate Earth-mass planet around a Sun-like star is disfavoured by every major method at once.

Coverage gaps not integrated in this release

Gaia DR3 is now cross-matched by exact source_id for every host the archive links to one — 4,408 systems — as an independent distance check, not a full astrometric re-reduction. The ESO Science Archive was investigated but is not built into this pipeline; see Data sources for the reasoning. Transit and RV analyses depend on public data existing at MAST and DACE respectively — most catalogue planets have neither, and this is reported as an explicit absence rather than omitted.

Gaia cross-check, in full

4,408 hosts matched. Median distance disagreement: 0.73%. 412 hosts flagged RUWE > 1.4 (possible unresolved binary).

Model and fit caveats

  • Transit-fit depths from this pipeline’s own trapezoid model run 5–25% below published values (a known detrending systematic) and are never substituted for catalogue values.
  • Radial-velocity semi-amplitude fits are gated by a three-criterion reliability check; a mass is withheld entirely when the fit fails it.
  • Habitable-zone membership from Monte Carlo draws is conditional on the draw’s temperature falling inside the model’s validity range — the fraction of draws that do is reported alongside the probability.
  • A high Earth Similarity Index alongside low observational confidence should never be read the same as a high index with high confidence; the two are reported as separate axes for exactly this reason.