The signal is the story you are trying to hear.
A signal is any repeatable or physically meaningful structure in the data. In astronomy, it might be a spectral line, a transit dip, a radial-velocity wobble, a pulsar timing pattern, a gravitational-wave chirp, or a faint source in an image.
The important word is not only "repeatable"; it is also "physically meaningful." A signal has a shape that should make sense before the plot seduces you. A planet transit should have a duration, depth, ingress, and egress that fit an orbit. A spectral line should land at a plausible wavelength after redshift and calibration. A pulsar should keep time with a precision that embarrasses ordinary clocks.
The measured data \(y(t)\) contain the true signal \(s(t)\) plus noise \(n(t)\). Real life adds systematics, calibration drift, sampling gaps, detector quirks, and deadlines.
A more honest version writes the model as a set of parameters. The data are not just "signal plus fuzz"; they are a measurement of something through an instrument.
Here \(m(x_i;\theta)\) is a model evaluated at observation \(x_i\), \(\theta\) is the parameter set, and \(\epsilon_i\) is what remains. The residuals are where confidence grows teeth or politely collapses.
Noise is everything that makes the plot less obedient.
Noise can come from photon statistics, electronics, background light, imperfect calibration, atmospheric effects, detector temperature, readout patterns, cosmic rays, sampling choices, or the star itself. Treating all noise as simple random scatter is how you accidentally write fiction with error bars.
The friendliest kind of noise is independent and random. If you repeat the measurement, it tends to average down. The mean becomes more stable, the uncertainty shrinks, and the universe starts sounding less like static. The less friendly kind is correlated or systematic: it has memory, structure, direction, and a suspicious talent for looking like the signal you wanted.
If you count \(N\) photons, the Poisson uncertainty is roughly \(\sqrt{N}\). That is why collecting four times as many photons improves SNR by about a factor of two, not four. Nature charges interest.
In precision astronomy, noise is not a single villain. It is a budget. Shot noise, read noise, dark current, sky background, flat-field errors, and calibration residuals all get a line item. If you do not know which term dominates, you do not yet know what improvement would actually help.
Smoothing is useful. It is also how people accidentally invent planets.
A smoothing window can reveal broad trends by suppressing high-frequency noise. But too much smoothing can erase real features, shift peaks, broaden dips, or create the illusion of structure. A beautiful curve is not automatically a truthful curve. Annoying, but important.
Smoothing is a filter. It decides which scales in the data are allowed to survive. That is not morally bad; every instrument has a resolution limit anyway. The danger begins when the smoothing scale is chosen after staring at the plot until it looks convincing.
A moving average replaces each point with nearby points. It reduces jaggedness, but it also makes adjacent points correlated. The smoothed curve has fewer independent pieces of information than it visually appears to have.
In this simplified form, the signal strength is compared to the spread of the noise. Different fields define SNR in more specific ways depending on the measurement, the model, and the statistics of the residuals.
A false positive is a discovery wearing a fake moustache.
The danger is not only missing real signals. It is believing in signals that are not there. That is why independent checks, null tests, injection-recovery tests, and physical plausibility matter.
If the detection disappears when you change a reasonable analysis choice, it was probably not the universe revealing itself. It was your pipeline doing interpretive dance.
False positives are especially good at hiding inside large searches. If you inspect one light curve, one spectrum, or one image, a rare fluctuation is rare. If you inspect a million windows, channels, targets, trial periods, and parameter combinations, a rare fluctuation starts acting like it paid rent.
If one test has false-alarm probability \(p\), then \(M\) independent trials raise the chance that at least one of them produces a tempting accident. Big searches need stricter standards.
A threshold is not a magic wand. It is a negotiated truce with uncertainty.
Scientists love thresholds because they turn messy evidence into a sentence: detected or not detected. But the threshold is not the discovery. It is a rule for deciding how much risk you are willing to tolerate.
A high threshold reduces false alarms but misses faint real signals. A low threshold catches more real signals but invites more impostors. The right choice depends on the cost of being wrong, the size of the search, the prior plausibility, and whether follow-up observations are possible.
The \(z\)-score says how many standard deviations a measurement \(x\) is from an expected mean \(\mu\). It is useful only when the noise model is honest. Non-Gaussian tails, correlated residuals, or underestimated uncertainties can make a glamorous \(z\)-score very fragile.
This is why astronomy papers often separate "candidate" from "confirmed." A candidate is interesting enough to deserve attention. A confirmed result has survived enough independent pressure that alternative explanations become less convincing.
A discovery should survive being treated with suspicion.
The mature version of data analysis is not cynicism. It is structured doubt. You want to give a real signal every fair chance to reveal itself while giving every fake signal several excellent opportunities to embarrass you privately before publication.
That workflow is slower than a dramatic claim. It is also how a dramatic claim becomes durable. The goal is not to make every result boring; the goal is to make the excitement expensive enough that only the sturdy signals can afford it.
Good science is not noise-free. It is noise-aware.
Data analysis is the discipline of being excited and suspicious at the same time. That tension is not a weakness. It is the reason discoveries survive contact with reality.
The cleanest plots are not always the most honest plots. The honest plot tells you what was measured, what was assumed, how uncertainty was handled, and what would make the claim fail. That is the kind of plot a discovery can stand on.