What a track record can and cannot tell you

A blind audit of four live trading track records — and of the instrument used to audit them.

About 3,400 trades across a window of roughly two and a half years ending in 2026. The subjects are anonymised as A, B, C and D, and the figures below are deliberately rounded: exact trade counts, exact timestamps and exact cash amounts are together identifying, and none of the findings needs that precision. This study is about what the method found, not about naming a commercial product.


1. What was found#

Subject A Subject B Subject C Subject D
trades (approx.) ~530 ~860 ~510 ~1,500
period ~6 months ~22 months ~14 months ~28 months
instruments 12 1 20 1
win rate ~69% ~72% ~61% ~77%
profit factor ~2.4 ~2.3 ~1.4 ~2.7
position sizing warn pass pass warn
true drawdown incl. open positions FAIL 11× warn 1.3× FAIL 2.4× FAIL 30×
stop-loss discipline not assessable not assessable not assessable not assessable
cost headroom warn (gated) pass ~30× warn (gated) warn (gated)
statistical assessability not assessable pass not assessable withheld — see §2
prop-firm pass probability not assessable pass ~87% not assessable withheld — see §2
capital injected into drawdown pass pass pass FAIL
exit-timing dependence pass pass pass pass

Three of the four fail on the geometry of their drawdown, and one fails on where its money came from. Not one fails on whether it made money.

The headline number is not the risk#

Three subjects report a true peak drawdown far larger than their account statement would suggest, because the loss sat in open positions that had not been closed yet:

drawdown a statement shows drawdown actually run ratio
Subject A ~170 ~1,800 11×
Subject C ~160 ~380 2.4×
Subject D ~80 ~2,300 30×

Subject D closed roughly fifteen hundred trades over more than two years, and its realised losing streak never cost it more than about eighty currency units. Its open positions, marked to an independent price feed, were some two thousand three hundred down at the worst point. A buyer reading the equity curve sees the first number. The second is the one that ends the account.

The deposits say what the equity curve does not#

Subject D's account received three equal deposits within twenty-six minutes of each other, late one evening — the last of them in the same minute as the peak of its worst floating drawdown. Before the next morning, a withdrawal larger than the account's entire trading profit had left it.

Across its whole record, more money was cycled through the account than the account ever earned: roughly thirteen thousand paid in and roughly sixteen and a half thousand taken out, against under five thousand of trading profit. Money arriving at the bottom of a drawdown is capital holding a losing position open, not capital compounding a winning one — and it converts a loss the equity curve would otherwise have had to show into fresh deposits.

This check exists only because these are live records. A backtest has one deposit and no owner.

One subject passed almost everything — so we built a check to try to break it#

Subject B is the only one with no failure. That combination — the cleanest report going to the record we trusted least — is the worst thing an audit can do, so the suspicion was turned into a test rather than left as a remark.

A live fill lands on an exact minute boundary about 1.7% of the time — one second in sixty. Across the four subjects:

entries on the minute exits on the minute
Subject A ~8% ~2%
Subject B ~3% ~15%
Subject C ~2% ~2%
Subject D ~88% ~77%

Subject C sitting on the 1-in-60 baseline in both columns is the control: it is what an untouched live tick record looks like, and it confirms the measurement reads the subjects rather than an artefact of the parser.

But clock-pinned exits are not a defect. Plenty of sound strategies close on a schedule, and a check that fired on the behaviour alone would be a timestamp histogram wearing a check's name. What matters is dependence: a follower — a copy-trader, a slower VPS, anyone reading a signal seconds late — never gets the exact instant the record claims. So the graded question is the one a buyer actually faces: does the reported profit survive those exits happening one minute later?

And a single delay is a verdict, where a sweep is a specification. So the delay is swept, and what is reported is the breaking point — the latency at which the record stops keeping half its profit. "Needs execution faster than X" is something a buyer can check their own setup against; "survives one minute" is not.

Re-priced against an independent minute-by-minute feed:

profit still standing after a delay of 60s 120s 300s 600s 30 min 60 min breaking point
Subject B (~15% pinned, ~130 exits) 99% 99% 99% 96% 98% 96% none to an hour
Subject D (~77% pinned, ~1,100 exits) 103% 100% 105% 116% 142% 124% none to an hour

The suspicion did not survive the test. Subject B's timing is a habit, not its result — an hour of delay still leaves 96% of its profit. Subject D's delay would have earned it more at every step. We say so as plainly as we would have reported the opposite, and the check keeps its place because on synthetic records it separates a timing-dependent result from an identically-pinned benign one.

One number does separate them, and it is a bound rather than a measurement. Our reference feed resolves to one minute, so a delay shorter than that cannot be priced — but it can be bounded, because a fill landing anywhere inside the exit minute gets a price within that minute's own high and low. Taking the worst instant in every exit bar:

profit surviving the worst instant inside its own exit minute
Subject B ~95%
Subject D ~51%

Subject D keeps barely half its profit if its exits land anywhere but the best moment of the minute they were booked in. That is a bound, not a forecast — the true figure sits somewhere between it and 103% — but the gap between the two subjects is the finding: B's result is robust inside the minute and D's is not. Where a feed carries no high and low, this bound is refused outright rather than quoted from closing prices, which would report a reassuring 100% for a record nobody had measured.

What remains unexplained is the behaviour, not the risk. Subject D's record is bar-quantised: an algorithm acting only on bar opens explains quantised entries, but not three-quarters of its exits — a market moving through a stop or a target does not wait for the minute. A scheduled close and a platform that rounds its timestamps are indistinguishable in this field. It is a question for the seller, not a finding against them.

Subject B closes on a timer rather than on a stop, which is worth asking a seller about even though it is not, on this evidence, worth holding against them. Neither observation is visible in any headline a buyer is shown.


2. What the audit found wrong with itself#

A study that reports only its subjects' failures is less credible than one that reports its own. Pointing this battery at live third-party data exposed three defects in the instrument; building the fixtures that let it ship exposed three more. All six had been invisible against simulated data, and all six ran in the subjects' favour — which is the direction they would run, because the records that break an audit are the unusual ones and unusual is where the losses are.

A seventh is in §4, and it is the only one that ran the other way: a field that looked like the missing stop-loss data, was not, and would have produced a false failure against a record. It is kept out of this section deliberately, because it is best read beside the check it would have broken.

1. The drawdown check passed an account on a withdrawal. Run as written, it reported Subject D's largest balance decline as a figure that matched one of its withdrawals to the currency unit — which was not a trading loss at all, but the owner moving money out — and, seeing the account's mark-to-market sitting barely 1% above that, concluded there was no meaningful open-position exposure. It issued a pass. With deposits and withdrawals removed from the curve, the same subject reports 30× and fails. A backtest has no cash movements, so this could not surface until a real account arrived. The error ran in the subject's favour.

2. A statistical check printed a verdict over a missing number. The frequency-ceiling calculation returns NaN when handed a zero trading cost, which is what it was handed. The check went on to print "not assessable against an N_max of nan" — which reads as a finding and was not one.

3. No check knew that another check's verdict was a precondition for its own. The two statistical checks consume daily returns as if they were independent draws. The drawdown check exists precisely to detect the case where they are not: a strategy that never closes a loser produces a smooth return series because the risk has not been realised yet. On Subject D these ran side by side and disagreed — Subject D's drawdown failing at 30×, while the statistical checks returned a clean bill of health for that same subject and put its prop-firm pass probability at about 70% — and nothing reconciled them. (Subject B's 87% in §1 is a different subject and a figure that still stands; the two are easy to run together and should not be.)

That last one was found by accident, on one subject. So every check was re-read against a single question — what does this check assume, and is there another check whose job is to detect that assumption failing? — and the dependencies are now declared and enforced, with two severities:

Subject D's two statistical passes are withheld under this rule. Before the gate they read a Sharpe of about 3.5 and a prop-firm pass probability near 70%. They are the numbers a vendor would quote, they were computed correctly, and they are not trustworthy — because the same file fails the check that tests their premise.

Three more, found by making every check prove itself#

No check may report a verdict here unless it has been shown to fire on a planted defect and stay silent on a control matched to it in every property but that defect. Seven checks had no such pair. Writing them found three more faults:

4. The drawdown check failed a record that had no drawdown at all. Its inflation figure is the ratio of true drawdown to the drawdown a statement shows, and a zero denominator was mapped to infinity — correct when open losses are what the statement hides, nonsense when there is nothing to hide. A position closed before the market moved scored: drawdown 0.00, statement drawdown 0.00, inflation infinite, fail. Found by the very control the check needed in order to ship.

5. Twelve places in the code returned a verdict with no direction of error and no statement of what it could not see — including the one printed most often in the whole product, because it appears on every record in this format. The paths that skipped them were the inconclusive ones, which is exactly where a reader most needs to know which way the missing answer would have cut. A finding can no longer be created at all without both.

6. The sub-minute bound in §1 quietly reported a reassuring 100%. It reads the high and low of the minute a trade exited in; the price feed was being loaded without them, so the bound computed as "no adverse move was possible". With the range loaded, Subject D's figure is 51%, not 100%. Where a feed carries no high and low, the bound is now refused outright rather than quoted from closing prices.

Two of these three would have made a report more flattering to a subject, and the third would have made one less legible. None was found by inspection. All three were found by requiring the instrument to demonstrate, on data with a known answer, that it does what it claims.

A fourth defect was predicted in writing before the run, and confirmed: the position-sizing check conditioned only on closed losses, so it could not see size added to a position still open. It now has a second arm that can. On a synthetic grid the new arm fires at 42×, where the old one saw nothing at all.

Two further corrections came from auditing the fix rather than the subjects:

The dependency rule was too blunt for a precondition that can never be met. Stop-loss discipline cannot be assessed on this data format at all, so making it a precondition risked permanently retiring the checks that depend on it — protection that amounts to retirement. A precondition that is unverifiable on a format is now recorded separately from one that merely did not run, becomes a standing qualifier on the dependent figure rather than a per-report caveat, and is explicitly not allowed to downgrade anything. Measured across the four subjects, every downgrade that did occur came from a genuine failure, never from an unverified precondition.

The exit-timing check's own thresholds were wrong on first writing. Its initial failure cut fired only when a delay removed all of a record's profit, which left a record keeping 1% of its profit reporting as a mild warning. The band was corrected against synthetic fixtures — before the check was ever run on a subject — to fail when more than half the profit is timing-dependent.


3. A correction this study makes to itself#

An earlier draft described two subjects as "opening at 15× and 25× the size of a losing position". Those figures are real but they are the maxima. The medians are 1.00× on all four subjects: these strategies add to losing positions at constant size.

That is averaging into a loser — a real exposure, and the reason three of them fail on drawdown — but it is not martingale escalation, and describing it as such overstates the case. Fewer than one adverse addition in ten exceeds 1.5×. Quoting a maximum as though it were typical is the same class of error this method exists to catch, pointed the other way.

The repaired check, run on the four subjects, changes no verdict.


4. What this data format could not support#

This is part of the finding, not a caveat attached to it. An audit that does not say what it could not see is not an audit.

check status on all four why
Stop-loss presence and distance never assessable no stop, target, close reason or comment field. Nothing distinguishes a stop-out from a discretionary close, and a stop that exists but never fires is invisible either way. A second format appeared to carry this field and did not — see below, because that is the more useful finding
Re-costing at a different spread not applicable this check's premise is a simulated record re-run at an assumed spread. These are live fills; the spread was already paid. Adding it back would charge it twice. Replaced by a cost-headroom measure that asks the live question instead
Bad-print exposure not assessable needs the vendor's own price feed; the export carries fills, not quotes
Financing convention not assessable the subjects' brokers are unknown
Look-ahead / timestamp integrity not assessable a closed-trade log shows fills, not what was known when the decision was made
Weekend and gap handling not assessable needs the vendor's own bars
Effective trial count not assessable a published signal exposes neither its inputs nor its optimisation history

And two checks were withheld on two subjects for a reason worth stating on its own: the reconstructed accounts implied leverage above 1:20,000 and above 1:7,000 respectively. No venue permits that, so the fault is not the trading — it is that the export's cash rows never reach the account's true opening balance. Every percentage-denominated figure would have been inflated by a denominator the file does not vouch for, so they were refused rather than scaled.

The stop-loss field that existed and was fake#

Not part of the sample. The four subjects above are the study. Separately, a record in a different export format was put through the same pipeline — to test what happens when an unfamiliar file arrives, not to audit it. It is not a fifth subject and nothing in §1 includes it. But what it showed about stop-loss discipline belongs here, because it is worth more to a buyer than the row in the table above.

That format carries a Comment column, and on 1,855 of 1,856 rows it holds exactly what the other format never has: an annotation naming how the trade closed, in the form [tp 1.35223]. On its face, the missing field, present at last. The check ran, and returned a clean, quotable failure: no exit in 1,855 trades was closed by a stop — this strategy carries no protective stop.

It was wrong, and it was wrong in the most persuasive way available. Two measurements settle it:

test result what it means
does the annotated price equal the close price? 1,855 of 1,855 it names where the trade ended, not why
does it sit on the losing side of the entry? 606 of 1,855 (33%) a take-profit cannot be below a Buy's entry

The field is the exit price wearing a tp label, applied whatever happened — including on 606 losing trades. Fed to the check, a naming convention became a finding.

And it corroborated. That record's drawdown check had already failed, on an account that closed no losers and carried its losses in open positions. A stop-loss failure sitting beside it read as two independent checks agreeing on one story. They were not independent and one of them was measuring nothing. A false finding that contradicts your other findings gets caught. One that confirms them does not.

The check now measures both properties before it will read the field at all, and declines when they fail. So stop-loss discipline remains unavailable on every format we read — but that is now a measured statement rather than an assumed one, and the availability of the check is a property of each record rather than a list in our code. The first export carrying a genuine close reason unlocks it with no change to the check.

For a buyer the transferable point is narrower and more useful than "the data was missing": a field being present is not the same as a field being real, and the difference is only visible if someone tests it against the prices.

The single field that would unlock the most#

An opening balance, or a balance series. It would restore statistical assessability on two of the four subjects immediately and make percentage drawdown defined on all four — currently it is refused everywhere, because net deposited capital varied by a factor of four in the mildest of the four records and by four orders of magnitude in the most extreme.

Second is a close reason or stop field — a genuine one. As above, we now test any such field against the record's own prices before reading it, and one that fails those tests unlocks nothing, however complete it looks.


5. None of the four traded often enough for costs to be the problem#

There is a frequency ceiling above which trading costs swallow a strategy's own claimed edge, however good the edge is. It is worth knowing where a record sits relative to it, because a strategy above its own ceiling cannot be rescued by trading better — only by trading less.

All four of these records sit far below theirs. They traded roughly 450 to 1,150 times a year against ceilings in the thousands to the tens of thousands. The ceiling was computed correctly and it was not close on any of them.

That is a finding about these four subjects, and a useful one: it rules out the explanation a buyer might reach for first. These records are not failing because they trade too much and pay too much to do it. Their problems are elsewhere — in the shape of their drawdown and in where their capital came from — and the cost check is what allows us to say so rather than guess.

It also sets expectations for this check's role. It was built against high-frequency candidates, where it decides the answer. On records that trade a few hundred times a year it will usually clear, and clearing is informative without being the headline.


6. Nine predictions, registered in advance: five right, four wrong#

All nine were written down and committed to version control before the battery was run against a single record. Five held. Four did not, and the four that did not are the more interesting half:

prediction outcome
Stops would be unassessable on all four right
Re-costing would not apply to any of them right
At least two would fail on true drawdown right — three did
No subject would pass everything right
The sizing check would under-detect open baskets right, and it did
Sizing: one outright failure, led by Subject A wrong — none failed
None would clear the statistical assessability bar wrong — one did
All four would fall below a 30% prop-firm pass rate wrong — two cleared it before gating
At least two would show an edge smaller than their spread wrong — none did; edges ran 7× to 16× the spread

All four errors point the same way: we expected these four records to die on trading costs and on statistical power, and they did not. They failed instead on the shape of their drawdown and on where their capital came from. That is the opposite of what simulated strategies had taught us to expect — for these four. Whether it generalises is a question four records cannot answer, and we are not going to pretend otherwise; it is the reason the next study is a larger sample rather than a deeper one.


7. Scope, stated so the numbers can be argued with#