Methodology

How we test whether this works.

PeerStock scores how closely a company resembles another company at another moment in time. That is only useful if the resemblance means something, so we measure it against random selection — and publish where it stops working.

What the score measures

Every company is described by 14 metrics — eight from its financial statements (revenue growth and acceleration, gross and operating and net margin, margin trend, free-cash-flow growth, return on equity) and six from its price history (12-month return, 6-month momentum change, 200-day trend slope, distance from the 52-week high, maximum drawdown, annualised volatility).

Each metric is converted to a percentile rank against every other company at that same moment, never to a raw figure. A 40% gross margin means something different in 2022 than today; ranking against contemporaries removes that problem. Similarity is then the average closeness of those ranks across all 14 metrics.

That comparison is deliberately sector-agnostic: a company is ranked against the whole market, not against its own industry. This is why matches are often not the obvious peers — a company can resemble another in growth, margin trajectory and price behaviour while operating in an entirely different business. Finding those non-obvious analogues is the point; screening by sector is something you can do yourself, and does not need us.

Metrics neither company has count as a miss rather than being quietly dropped, so a company matching well on three metrics cannot outscore one matching well on fourteen.

No look-ahead bias

This is the constraint everything else is built around. If you search Apple as it looked on 24 May 2023, every figure used must have been public on 24 May 2023.

The trap is that a fiscal period ends months before its results are filed. A quarter ending in March may not be published until June — so filtering on the period end would let the engine read results that did not yet exist. PeerStock stores the actual filing date recorded on each SEC fact and filters on that instead. Not an estimate of the reporting delay: the date the document was filed.

An automated test suite walks the full provenance chain behind every metric on every change, and fails the build if any fact carries a filing date later than the search date. The rule is enforced mechanically rather than by remembering to be careful.

How the test is constructed

Take a company at a historical date — the reference. Run the search as it would have run at a later date, and record what the matches did over the following six months. Separately, record what the reference itself did over the six months after its date. The question is whether the matches behaved as the reference did.

Alongside each match’s own return we therefore record a distance: for each match, the gap between its own six-month result and the reference’s own six-month result, in percentage points. A match that gained 30% where the reference gained 34% misses by 4 points. This is deliberately not a measure of whether matches went up — a tool that simply picked stocks that rose would score badly here whenever the reference fell, and that is the correct outcome for a tool whose claim is resemblance.

The control is random stocks drawn from the same scored universe, over identical start and end dates. That matters more than it sounds: comparing matches against a benchmark from a different period would credit or blame the engine for whatever the market did in between. Same universe, same window, same conditions — the only difference is whether the similarity score chose them.

Results are also grouped by what the reference went on to do — gained 25% or more, stayed within ±5%, or fell 25% or more — and each group is compared only against its own control. A group of references that rose is drawn from periods when the market may have risen too, and a shared control would hand the engine credit for that. References landing between those bands are discarded rather than pushed into a neighbouring group, so “stable” means stable.

Results

Across 922 historical reference periods and 3.9 million scored candidates: when the reference company went on to gain 25% or more, the margin over random picks scaled with the match score. Random picks over those same dates returned −0.5%, so the return columns below are a comparison against chance rather than a return anyone earned.

A second measure follows below, and on it the effect is considerably larger: a 90%+ match finished 17.0 points nearer the reference’s own outcome than a random pick did, against a 2.9 point edge on raw return. Both are reported because they are the same 922 tests read two ways — and the smaller, plainer one is the one most readers can check for themselves.

Why 25%? We specifically tested historical periods where the reference company subsequently gained 25% or more, because these are the situations where users are most likely to want historical precedents. Periods were selected by what the reference company went on to do — never by what the matches did — and each is compared against a control drawn from those same periods, so a rising market cannot be mistaken for the engine working.
Match scoreCandidatesMedian 6-month returnvs. randomrandom: −0.5%
90%+942+2.3%+2.9 pts
85–89%16,951+1.2%+1.7 pts
80–84%79,918+0.7%+1.2 pts
75–79%183,293−0.2%+0.4 pts
below 75%3,616,694−0.4%+0.2 pts

The same tests, measured against the reference

The table above asks whether the matches went up. This one asks the question the engine actually answers: how close did each match’s six-month result land to the reference’s own six-month result? Lower is better, so the second column reports it as an improvement on the control.

CompanySnapshotIts own next 6 months
ReferenceAAPL28 Jun 2019+46.4%
Match90% similarSPXC30 Jun 2025+22.7%
Misshow far apart the two outcomes landed≈ 24 points

The table below pools 942 comparisons like that one — every 90%+ match from all 922 reference periods — and reports the median. Half miss by less than 53.0 points, half by more. The pair above is a good one, and a single pair proves nothing either way: what the table shows is that across all of them, a random pick measured the same way typically misses by 70.0.

Match scoreCandidatesMiss vs. referenceCloser than randomrandom missed by 70.0%
90%+94253.0%+17.0 pts
85–89%16,95154.1%+15.9 pts
80–84%79,91856.4%+13.7 pts
75–79%183,29359.5%+10.5 pts
below 75%3,616,69470.9%−0.9 pts

The unbroken ordering is the finding, more than any single figure. One bucket doing well can happen by chance; every band improving step by step, in the same order as the score, on both measures at once, is considerably harder to produce by accident. It indicates the score measures degree of resemblance, not merely its presence.

That ordering needs a great many reference periods to be trustworthy, which is the reason for 922 of them. Every candidate inside a single search shares that search’s reference, so the meaningful sample size is the number of reference periods, not the 3.9 million candidates. Tested on roughly a hundred, the ordering moves around between runs; at this size it holds.

What these results do not show

This is not a prediction of returns. It measures whether companies that look alike went on to behave alike. A high score is a reason to look closely at a company, not a reason to buy it.

Data and coverage

See it on a company you know

The fastest way to judge any of this is to search a company whose history you already understand, and read the evidence behind the matches.

Try it now

Research tool · not investment advice · PeerStock does not recommend securities.