PeerStock scores how closely a company resembles another company at another moment in time. That is only useful if the resemblance means something, so we measure it against random selection — and publish where it stops working.
Every company is described by 14 metrics — eight from its financial statements (revenue growth and acceleration, gross and operating and net margin, margin trend, free-cash-flow growth, return on equity) and six from its price history (12-month return, 6-month momentum change, 200-day trend slope, distance from the 52-week high, maximum drawdown, annualised volatility).
Each metric is converted to a percentile rank against every other company at that same moment, never to a raw figure. A 40% gross margin means something different in 2022 than today; ranking against contemporaries removes that problem. Similarity is then the average closeness of those ranks across all 14 metrics.
That comparison is deliberately sector-agnostic: a company is ranked against the whole market, not against its own industry. This is why matches are often not the obvious peers — a company can resemble another in growth, margin trajectory and price behaviour while operating in an entirely different business. Finding those non-obvious analogues is the point; screening by sector is something you can do yourself, and does not need us.
Metrics neither company has count as a miss rather than being quietly dropped, so a company matching well on three metrics cannot outscore one matching well on fourteen.
This is the constraint everything else is built around. If you search Apple as it looked on 24 May 2023, every figure used must have been public on 24 May 2023.
The trap is that a fiscal period ends months before its results are filed. A quarter ending in March may not be published until June — so filtering on the period end would let the engine read results that did not yet exist. PeerStock stores the actual filing date recorded on each SEC fact and filters on that instead. Not an estimate of the reporting delay: the date the document was filed.
An automated test suite walks the full provenance chain behind every metric on every change, and fails the build if any fact carries a filing date later than the search date. The rule is enforced mechanically rather than by remembering to be careful.
Take a company at a historical date — the reference. Run the search as it would have run at a later date, and record what the matches did over the following six months. Separately, record what the reference itself did over the six months after its date. The question is whether the matches behaved as the reference did.
Alongside each match’s own return we therefore record a distance: for each match, the gap between its own six-month result and the reference’s own six-month result, in percentage points. A match that gained 30% where the reference gained 34% misses by 4 points. This is deliberately not a measure of whether matches went up — a tool that simply picked stocks that rose would score badly here whenever the reference fell, and that is the correct outcome for a tool whose claim is resemblance.
The control is random stocks drawn from the same scored universe, over identical start and end dates. That matters more than it sounds: comparing matches against a benchmark from a different period would credit or blame the engine for whatever the market did in between. Same universe, same window, same conditions — the only difference is whether the similarity score chose them.
Results are also grouped by what the reference went on to do — gained 25% or more, stayed within ±5%, or fell 25% or more — and each group is compared only against its own control. A group of references that rose is drawn from periods when the market may have risen too, and a shared control would hand the engine credit for that. References landing between those bands are discarded rather than pushed into a neighbouring group, so “stable” means stable.
Across 922 historical reference periods and 3.9 million scored candidates: when the reference company went on to gain 25% or more, the margin over random picks scaled with the match score. Random picks over those same dates returned −0.5%, so the return columns below are a comparison against chance rather than a return anyone earned.
A second measure follows below, and on it the effect is considerably larger: a 90%+ match finished 17.0 points nearer the reference’s own outcome than a random pick did, against a 2.9 point edge on raw return. Both are reported because they are the same 922 tests read two ways — and the smaller, plainer one is the one most readers can check for themselves.
| Match score | Candidates | Median 6-month return | vs. randomrandom: −0.5% |
|---|---|---|---|
| 90%+ | 942 | +2.3% | +2.9 pts |
| 85–89% | 16,951 | +1.2% | +1.7 pts |
| 80–84% | 79,918 | +0.7% | +1.2 pts |
| 75–79% | 183,293 | −0.2% | +0.4 pts |
| below 75% | 3,616,694 | −0.4% | +0.2 pts |
The table above asks whether the matches went up. This one asks the question the engine actually answers: how close did each match’s six-month result land to the reference’s own six-month result? Lower is better, so the second column reports it as an improvement on the control.
| Company | Snapshot | Its own next 6 months | |
|---|---|---|---|
| Reference | AAPL | 28 Jun 2019 | +46.4% |
| Match90% similar | SPXC | 30 Jun 2025 | +22.7% |
| Miss | how far apart the two outcomes landed | ≈ 24 points | |
The table below pools 942 comparisons like that one — every 90%+ match from all 922 reference periods — and reports the median. Half miss by less than 53.0 points, half by more. The pair above is a good one, and a single pair proves nothing either way: what the table shows is that across all of them, a random pick measured the same way typically misses by 70.0.
| Match score | Candidates | Miss vs. reference | Closer than randomrandom missed by 70.0% |
|---|---|---|---|
| 90%+ | 942 | 53.0% | +17.0 pts |
| 85–89% | 16,951 | 54.1% | +15.9 pts |
| 80–84% | 79,918 | 56.4% | +13.7 pts |
| 75–79% | 183,293 | 59.5% | +10.5 pts |
| below 75% | 3,616,694 | 70.9% | −0.9 pts |
The unbroken ordering is the finding, more than any single figure. One bucket doing well can happen by chance; every band improving step by step, in the same order as the score, on both measures at once, is considerably harder to produce by accident. It indicates the score measures degree of resemblance, not merely its presence.
That ordering needs a great many reference periods to be trustworthy, which is the reason for 922 of them. Every candidate inside a single search shares that search’s reference, so the meaningful sample size is the number of reference periods, not the 3.9 million candidates. Tested on roughly a hundred, the ordering moves around between runs; at this size it holds.
The fastest way to judge any of this is to search a company whose history you already understand, and read the evidence behind the matches.
Try it nowResearch tool · not investment advice · PeerStock does not recommend securities.