A reviewer-facing record of the Top-50 repository upgrade programme. Each completed entry connects a research question to generated results, validation, a versioned release, and an explicit boundary of inference.
121repositories audited
115eligible first-party
18 / 50upgrades complete
32ranked upgrades remaining
Completion rule. A repository is complete only after a scoped research question, generated evidence, explicit limitations, automated validation, a public release, and live-link verification.
Queue #1
EPRV Spectrograph Landscape
v0.15.0 · verified
Research questionHow strongly does resolving power improve the local photon-limited Doppler bound as intrinsic stellar line width changes?
Generated result
Increasing resolving power from 100,000 to 150,000 improves the controlled bound by 37.1%, 11.1%, and 2.48% for intrinsic Gaussian widths of 1.0, 2.5, and 5.0 km/s. The broad-line null below 10% is not rejected.
Evidence
15 controlled scenarios
4,000 seeded recoveries
27 tests
Python 3.10/3.12/3.13 CI
Evidence maturity
Before63
After92
+29 rubric points
Queue #4
ExoLight Transit Lab
v1.4.0 · verified
Research questionHow badly can correlated time-series noise invalidate iid transit-depth uncertainties, and can event-clustered intervals restore coverage?
Generated result
With AR(1) correlation ρ=0.8, nominal 95% iid coverage falls to 54.9% and 56.1% for deep and shallow injections; event-clustered intervals reach 93.2% and 92.1%.
Evidence
4 controlled scenarios
4,000 seeded recoveries
deterministic reports
Scientific core, frontend, and Pages CI
Evidence maturity
Before58
After95
+37 rubric points
Queue #5
EXOhSPEC Th–Ar Atlas
v0.3.0 · verified
Research questionCan a multi-exposure ladder preserve strong calibration-feature centroids better than a single long exposure without degrading faint features?
Generated result
Across 7,000 strong-feature injections, pixelwise HDR lowers centroid RMSE from 0.012954 to 0.002086 pixel, an 83.9% reduction; faint-feature RMSE remains unchanged.
Evidence
24 measured feature strengths
24,000 seeded recoveries
6 tests
Python 3.10/3.12/3.13 CI
Evidence maturity
Before52
After93
+41 rubric points
Queue #6
Research Portfolio Evidence Registry
v2026.09.24.3 · verified
Research questionCan a large research portfolio expose questions, generated results, validation, releases, and limitations in one reviewer-auditable layer?
Generated result
The first sixteen completed upgrades are represented by one validated schema and a deterministic public registry, while the page states that 34 ranked repositories remain unfinished.
Evidence
121 repositories audited
50-repository ranked queue
schema validation
deterministic HTML build and CI
Evidence maturity
Before60
After95
+35 rubric points
Queue #7
LATTICE Radiation Twin
v0.2.0 · verified
Research questionCan eight historical HST epochs distinguish calendar time from four cumulative radiation proxy components well enough to interpret individual coefficients?
Generated result
The component-attribution gate fails: six parameters leave two residual degrees of freedom, the standardized condition number is 175.31, maximum VIF is 4,881.4, and three exposure coefficients change sign under leave-one-epoch-out refits.
Evidence
8 independent epochs
4 identifiability diagnostics
82 tests
Scientific and Pages CI
Evidence maturity
Before83
After98
+15 rubric points
Queue #8
Finding Earth 2.0
v2.1.2 · verified
Research questionWhen the scientific objective is reducing planet-radius uncertainty, can a larger scalar stellar-radius information score justify an indirect observation without measured covariance?
Generated result
Thirteen targets support both actions. Stellar radius has the larger own-parameter score for 10, but the indirect route wins for none at |rho| <= 0.90; Kepler-296 f requires |rho| = 0.995225 to break even.
Research questionHow sensitive is a 0.05 AU and 140 m close-approach screen to JPL distance intervals, missing diameters, assumed albedo, and row-limited query coverage?
Generated result
The reviewed 7,500-row chronological prefix contains 586 nominal-close approaches, of which 396 remain close at the three-sigma maximum distance; 190 nominal-close rows cross the distance boundary. Only 235 rows have measured diameters, and 61 nominal-close rows are albedo-sensitive around the 140 m screen.
Evidence
7,500 of 39,427 matching JPL CAD rows with explicit 19.0% coverage
235 measured diameters and 7,263 H-derived sensitivity intervals
Six numerical regression tests and deterministic JSON/CSV/SVG products
Node 24 validation, static build, Pages gate, and zero npm vulnerabilities
Evidence maturity
Before39
After96
+57 rubric points
Queue #10
GitHub Profile Evidence Index
v1.1.0 · verified
Research questionCan a GitHub profile distinguish reviewer-auditable research evidence from general portfolio discovery while keeping every concise claim traceable to a versioned source?
Generated result
Seven upstream records now route each research question and generated result to an exact source revision, versioned release, live evidence surface, and project-specific boundary; discovery-only links are explicitly excluded from that verified set.
Evidence
Seven pinned upstream evidence records
Deterministic JSON-to-Markdown/CSV/SVG build
Five schema, freshness, traceability, uniqueness, and boundary tests
Node 24 CI and zero dependency vulnerabilities
All 21 repository/release/live evidence URLs verified HTTP 200
Four release assets with GitHub-verified SHA-256 digests
Evidence maturity
Before36
After95
+59 rubric points
Queue #11
Spectral Noise Budget Studio
v1.1.0 · verified
Research questionAt four declared molecular feature wavelengths, which committed JWST/NIRCam profiles retain at least half of peak response?
Generated result
Across 16 feature/filter comparisons, four are in the half-power core, one is an edge placement, and eleven are outside useful sampled response. CH₄ at 3.30 μm in F277W retains about 0.026% of peak response, a 3,805× source-photon-limited relative-time proxy.
Evidence
3,192 committed SVO throughput samples with source SHA-256 receipts
16-row deterministic JSON and CSV audit
Linear interpolation, pivot wavelength, and equivalent-width calculations
13 data-contract, numerical-property, and headline-result tests
Node 24 CI, Pages deployment, and zero dependency vulnerabilities
Four release assets with GitHub-verified SHA-256 digests
Evidence maturity
Before44
After95
+51 rubric points
Queue #12
Jana’s RV Doppler Observatory
v4.1.0 · verified
Research questionAre the 250 committed RV target files inference-ready under declared numerical, duplication, provenance, offset-separability, and outlier contracts, and can the browser analysis avoid instrument-zero-point confounding?
Generated result
Only 39 of 250 files pass all six inference-readiness gates. The audit finds 34,153 rows in repeated-epoch groups, 181 files with reference/offset confounding, and 44 files with extreme full-to-robust RV-span ratios; non-ready bundled files are now blocked from automatic scanning and fitting.
Evidence
154,150 committed rows and 250 canonical SHA-256 file receipts
Group-centered weighted candidate-period scan with no false-alarm claim
Analytic amplitude and per-instrument systemic-offset solution in a bounded fixed-period grid
Four v4.1 interface-evidence assets with GitHub-verified SHA-256 digests
Evidence maturity
Before32
After96
+64 rubric points
Queue #13
Adaptive Optics Wavefront Lab
v3.0.0 · verified
Research questionDoes the 300-frame browser reduction preserve the values, distributions, and temporal content of all 15,000 released CIAO adaptive-optics telemetry frames?
Generated result
Every retained sample matches the released FITS source within the declared six-decimal rounding, and all three distribution gates pass. The temporal-fidelity null is rejected: 20.89%, 9.19%, and 64.46% of full-rate gradient-RMS, command-RMS, and mean-flux power lies above the reduced product's 5 Hz Nyquist frequency.
Evidence
15,000 released CIAO1 frames audited against 300 browser samples
Exact retained-point parity within six-decimal rounding
Three distribution gates pass and the predeclared temporal-fidelity gate fails
Released FITS source DOI, programme metadata, and MD5 receipt
Node 24 and Python 3.12 research verification plus Pages deployment
Six release assets with GitHub-verified SHA-256 digests
Evidence maturity
Before51
After95
+44 rubric points
Queue #14
Synthetic Point-Lens Microlensing Recovery Study
v0.5.1 · verified
Research questionUnder a declared Roman-motivated cadence and white-noise proxy, does crossing a blind event-detection threshold imply that the injected t0, u0, and tE are identifiable?
Generated result
The blind search detects 864 of 900 balanced synthetic injections (96.0%, Wilson 95% CI 94.5–97.1%) but recovers the predeclared parameters in 710 (78.9%, 76.1–81.4%). At F146=24 and tE=0.02 d, detection is 41/75 while parameter recovery is 22/75; 0/1,000 simple constant-flux nulls trigger, constraining that scoped null rate below 0.383% at 95% confidence.
Evidence
900 injections with unique parameter-aware master, cadence, epoch, and noise seeds
1,000 identically searched null realizations with raw rows and Wilson intervals
Truth-blind matched-filter proposal followed by bounded fitting and full-epoch chi-square evaluation
39 tests, Ruff validation, and Python 3.10/3.12 CI
Failed custom binary-lens validation preserved as a negative result and runtime-disabled
Evidence maturity
Before39
After95
+56 rubric points
Queue #15
K2-18 b Spectral Robustness Audit
v1.0.0 · verified
Research questionDo the repository's descriptive 4.3 micrometre contrast and a diagonal-error MIRI flat-spectrum test remain significant under predeclared window and leave-one-bin-out perturbations?
Generated result
The historical window reproduces at 4.07 ± 23.16 ppm (0.176 sigma), while 30 predeclared window choices span -0.861 to +2.258 sigma and none reaches |3 sigma|. The 27-bin MIRI flat fit gives p = 0.000968, but deleting the 5.375 micrometre bin raises p to 0.07832, so the result fails the declared every-one-bin-deletion robustness rule.
Evidence
4,411-point NIRISS+NIRSpec spectrum and separate 28-bin MIRI source product with canonical SHA-256 receipts
30 fully reported band/continuum definitions
27 leave-one-bin-out MIRI refits
NASA Exoplanet Archive parameter snapshot with exact TAP query
17 tests, Ruff validation, and Python 3.10/3.12/3.13 CI
Deterministic numerical JSON/CSV evidence and semantically validated SVG figures
Evidence maturity
Before45
After96
+51 rubric points
Queue #16
WASP-80 b Native-Channel Morphology Audit
v1.0.0 · verified
Research questionDoes the descriptive residual contrast between the published full and component-removed CH4 curves keep the same sign across declared adjacent-channel bin widths and origins in both transmission and emission?
Generated result
All 113 native source-data channels were reconstructed. The component-removal residual contrast remains positive across all 30 declared designs per geometry: 63.17 to 85.74 in transmission and 217.65 to 245.99 in emission. The result is posterior morphology sensitivity, not an independent methane detection test.
Evidence
113 native JWST/NIRCam source-data channels with source and derivative SHA-256 receipts
30 declared bin-width/origin designs per geometry and 60 fully reported rows
Portable spectral evidence gates and tolerance-based nonlinear optimizer regression
Evidence maturity
Before50
After94
+44 rubric points
Queue #17
Spectrograph RV Proxy Generalization Audit
v2.0.0 · verified
Research questionAfter calibration at one fiducial spectrum, does a linear resolving-power proxy preserve the discrete photon-information bound across stellar line-width, depth, and line-count regimes?
Generated result
Across 288 predeclared synthetic spectra, the calibrated legacy proxy is at least twofold too optimistic in 39 scenarios and reaches a maximum exact-to-proxy uncertainty ratio of 2.924. From R=100,000 to 250,000 it claims a 60% gain for every line width, while the exact median gain falls to 5.86% for 10 km/s intrinsic lines and rises to 66.89% for 1 km/s lines.
Evidence
288 fully reported scenarios across resolving power, intrinsic width, depth, and line count
Discrete Bouchy-style Poisson Fisher bound with fixed coverage and continuum-electron budget
Four-versus-eight-pixel convergence with maximum relative difference 1.162e-6 and zero primary-classification changes
Nine numerical tests and Node 20/22/24 CI
Protocol and scientific-model SHA-256 bindings plus byte-stable generated evidence
Seven release assets with GitHub-verified SHA-256 digests
Evidence maturity
Before41
After95
+54 rubric points
Queue #18
WASP-39 b Supplied-Model Sensitivity Audit
v2.0.0 · verified
Research questionDoes the descriptive preference for the archived full ScCHIMERA curve over its remove-CO2 curve retain its sign under strict common support, model integration, asymmetric-error rules, every rebin origin, covariance stress tests, and contiguous wavelength-block deletion?
Generated result
All 60 declared preprocessing designs retain a positive supplied-model contrast, spanning delta chi-square 559.6 to 780.4, and every single-bin deletion remains above 684.8. Yet deleting a 0.5-micrometre block around the 4.3-micrometre feature reduces the contrast to 3.19 and the outside-4.1-to-4.6-micrometre refit gives 6.04, demonstrating preprocessing stability but strong spectral localization.
Evidence
93 strict common-support bins from the 94-bin Eureka PRISM spectrum
20-cell illustrative covariance grid with eigenvalue and condition-number diagnostics
All declared single-bin and 0.10/0.25/0.50-micrometre contiguous-block deletions
Cross-platform offline SHA-256 verification against exact Zenodo ZIP entries
14 tests and Python 3.11/3.12 CI with SHA-pinned Node 24 actions
Seven release assets with GitHub-verified SHA-256 digests
Evidence maturity
Before55
After95
+40 rubric points
Queue #19
WASP-43 b Phase-Curve Identifiability Audit
v2.0.0 · verified
Research questionDoes a pre-eclipse first-harmonic maximum remain directionally stable across reductions, wavelength windows, aggregation rules, uncertainty conventions, and wavelength deletion, and can four archived phase spectra identify the source paper's two-harmonic model?
Generated result
All 180 predeclared first-harmonic designs place the maximum before eclipse, spanning -9.71 to -7.66 degrees with a median of -8.82 degrees; leave-one-trusted-wavelength results span -9.72 to -8.95 degrees. Four phase observations identify the three-parameter first harmonic with one residual degree of freedom, while the five-parameter two-harmonic design has rank four and is underidentified.
Evidence
Two HDF5 spectra verified byte-for-byte against exact Zenodo archive members
11 trusted 5.25-10.25 micrometre bins used for inference and three shadowed-region bins retained only for provenance/display
22 leave-one-wavelength-out refits across both reductions
Explicit one- versus two-harmonic design-matrix rank audit
13 tests and Python 3.11/3.12 CI with SHA-pinned Node 24 actions
Seven v2.0.0 release assets with GitHub-verified digests
Evidence maturity
Before48
After94
+46 rubric points
Queue #20
HAT-P-11 b Helium Estimator Sensitivity Audit
v2.0.0 · verified
Research questionDoes the transit-centred helium dip remain positive across declared phase windows, baseline exclusions, baseline sides, weighting rules, and point deletion while preserving exact archive provenance?
Generated result
All 74 valid predeclared estimator designs retain a positive helium dip, spanning 0.511% to 1.160% with a median of 0.841%. The 16 default-design leave-one-point-out estimates span 0.747% to 1.003%; formal signal-to-noise values of 3.64 to 14.06 are conditional on independent reported errors and are not an independent detection significance.
Evidence
Two renamed inputs matched to exact Zenodo ZIP members by canonical-LF SHA-256
74 fully reported phase-window, exclusion, weighting, and baseline-side designs
16 leave-one-point-out estimates for the default in- and out-of-transit samples
Published 1.08 +/- 0.05% result kept separate from the repository estimator
Four tests and Python 3.11 CI with SHA-pinned Node 24 actions
Six v2.0.0 release assets with GitHub-verified SHA-256 digests
Evidence maturity
Before44
After94
+50 rubric points
What this registry does not claim
Maturity scores measure the presence of auditable research practices across ten documented dimensions. They are not peer-review scores, citation metrics, or literal multipliers of scientific quality. Repositories not shown here remain in the ranked queue and are not represented as upgraded.