01
The catalogue is shaped by the telescope
A list of discoveries is not a census of the Galaxy. Transit geometry, observing windows, noise, pipeline recovery and vetting decide which planets become visible to us.
I began with a simple question and built a reproducible research system around it. The deeper lesson was that finding another Earth is less about sorting a catalogue and more about learning exactly what each observation can—and cannot—tell us.
164,209 source records · 6,354 confirmed planets · every claim labelled by evidence type
Why I built it
I wanted to know which known world comes closest to Earth. At first that sounded like a ranking problem. Building the pipeline showed me that it is really an evidence problem: the catalogue mixes observations, derived quantities, model estimates and missing uncertainty.
So I built more than a score. I built a source trail, uncertainty propagation, Solar-System controls, observable-specific laboratories and a growing dictionary that forces me to name what kind of claim I am making.
The project now asks a better question: given how our surveys select what we see, what does the evidence imply about the population beyond the visible catalogue—and which observation would teach us the most next? Reconstructing the Kepler DR25 selection function made that lesson quantitative.
Learning log
01
A list of discoveries is not a census of the Galaxy. Transit geometry, observing windows, noise, pipeline recovery and vetting decide which planets become visible to us.
02
Masses and other parameters can be inferred, calculated by an archive, or copied through composite records. I learned to ask where every number came from before using it.
03
Size, irradiation, composition, atmosphere, observability and biology require different evidence. Combining them too early creates confidence the observations do not support.
04
A planet is better represented by a distribution than a point. Propagating published error bars exposes fragile ranks and prevents false precision.
05
Venus looks deceptively Earth-like to a bulk-property score. Running the Solar System through the same machinery reveals what the score cannot know.
06
A useful search should identify the measurement that would reduce uncertainty most, rather than ending with a static ranking.
Research dictionary
This dictionary is part of the method. Each label limits what a number is allowed to claim, from a direct observation to a conditional forecast.
Author perspective
These are my hypotheses, values and open questions. They are not NASA or ESA conclusions and they are not outputs of the candidate model.
Author roadmap · living project
The next version should earn more confidence, not add a more confident score. My plan is to shorten the path from a changing archive record to a reproducible decision about the next useful observation.
01 · CONTINUOUS
Archive updates → validated releases
Refresh public archives on schedule, preserve every source hash, and publish clear diffs when the catalogue or a conclusion changes.
02 · NEXT MODEL
Separate errors → joint decisions
Build joint stellar-and-planet posteriors, add instrument likelihoods and compare information gain per unit of observing time.
03 · NEXT EVIDENCE
Scenarios → testable evidence
Add target-specific ephemerides, XUV histories and retrieval evidence while keeping reductions and natural alternatives attached.
04 · LONG HORIZON
Forecasts → observations
Version adapters for Gaia DR4, PLATO, Roman, ELT and HWO products only when public data exist, then revisit distant-world and contact limits.
Automatic data pulse
The full archive pipeline downloads, validates, hashes and processes data in an automated workflow. Scientific changes remain review-gated before publication; this browser check can surface the newest committed result without pretending that an unchecked archive row is analysis-ready.
Checking the latest research snapshot…
Author conclusion
“I did not find a second Earth. I learned how difficult it is to earn that conclusion—and built a system that makes every step of the search inspectable.”
The most honest result is not a winner. It is a map of what humanity has measured, what the models add, where the selection effects hide, and what evidence is still missing. My plan is to keep that map alive, make its uncertainty more connected, and let new observations—not ambition alone—move its conclusions forward.