Testing Our Calculator, KeepTradeCut, FantasyCalc, and RosterAudit on the Same Trades
We benchmark our trade verdicts against how real dynasty communities judge real trades, and we've written about that test before. The obvious follow-up question is how other calculators score on the same test. So we ran it, against KeepTradeCut and FantasyCalc, the two tools dynasty managers cite most, and RosterAudit, a newer calculator that has been earning attention. The test covers 4,461 community-judged trades from the r/DynastyFF trade megathreads and r/DynastyFFTradeAdvice, October 2025 through August 2026. All of these are good free tools that we respect. KeepTradeCut's values are crowdsourced from millions of community votes, while FantasyCalc's, RosterAudit's, and ours are learned from real trades, each in a different way. This article shows the results and the work behind them.
The rules of a fair test
A calculator comparison is only fair if every tool is tested as it actually behaves. Judging a competitor on raw value totals alone would understate how it really performs, because each calculator layers on its own logic. KeepTradeCut applies a value adjustment and its own fairness rule. FantasyCalc applies a waiver adjustment. RosterAudit grades the gap between plain sums. We ran each tool's full real logic and verified the implementations by reproducing live results from each site exactly, across one-for-one swaps, two-for-ones, and three-for-ones, in both Superflex and one-QB formats.
“Even” itself has to be defined per tool, so each calculator got its own real rule where one exists. KeepTradeCut's site calls a trade fair when the adjusted totals differ by no more than 5% of their combined value, its default setting, and that is the rule we ran. RosterAudit's calculator calls a trade even when the gap is within 10% of the larger side, and that is the rule we ran for it. Our verdicts use our live calculator's rule, also even within 10% of the larger side. FantasyCalc's live calculator has no fair band at all and always names a winner, however thin the margin, so for the purposes of this exercise we gave it our same 10% band, applied after its waiver adjustment. The band-by-band table below runs every calculator at every band from 4% to 16%, so no conclusion depends on that assignment.
Every trade is priced with each system's values as of the day the thread was posted, with no hindsight, and the results come only from trades all the compared systems could fully price, so nobody gets graded on a trade it couldn't evaluate. An even verdict is still a verdict. When the community saw a clear winner and a calculator said even, that calculator got the trade wrong, just less loudly than picking the losing side. So we grade two things separately, how often each tool is right when it names a winner, and how often it handed you the community's answer at all.
How the community's answer is decided
The ground truth for this test is what each thread's replies actually said. Every distinct commenter is classified as backing one side, backing the other, or calling the trade fair, with hedged leans counted at reduced weight and the poster excluded from the count of his own trade. One person is one voice however many times they reply, weighted by the upvotes on their best-received comment. A thread earns a clear verdict only when a real majority of the weighted voices agrees and at least two distinct commenters are on the majority side. Threads with too few substantive replies to mean anything are excluded rather than guessed. Of the 4,461 judged threads, 3,890 produced a usable community verdict, and 3,004 of those could be fully priced by us, KeepTradeCut, and FantasyCalc together. RosterAudit is scored on a separate corpus below: its published values only begin in June 2026, and on the window they cover, 1,281 of these trades can be priced by all four systems at once. The full counting rules, with worked threads, are in How a Reddit Thread Becomes a Verdict.
One note on “fair” verdicts before the results. The community almost never issues them. Only 201 of the 3,004 trades, about 7%, ended with the replies genuinely split or calling the trade even, and no calculator's even calls line up with those trades much better than chance, ours included. So the scoring below focuses on the question every tool can be graded on, who won.
The full process on two real trades
Here are two real trades, worked through the same way every trade in the study was. Names arrive as the community writes them (“Kyren,” “Dobbins,” “Achane”) and are resolved deterministically against the player database, with ambiguous names resolved to the fantasy-relevant player and unresolvable ones dropped rather than guessed. Each tool then prices both sides from its own values for that exact date, and each tool's real verdict logic renders the call. The community's replies, not ours, decide who was right.
Trade 1, one-QB league, July 2026: Josh Jacobs and Kyren Williams for Omarion Hampton.
The replies backed the Hampton side in one voice: “Fleece job,” “You sold two declining assets for a young rb1,” “I'm taking Hampton all day.”
| Calculator | Jacobs + Williams | Hampton | Verdict |
|---|---|---|---|
| KeepTradeCut | 4,697 + 5,625 = 10,322 | 7,378 + 2,505 adj = 9,883 | Fair trade, 2.2% apart |
| FantasyCalc | 2,942 + 3,712 = 6,654 | 6,357 + 753 waiver = 7,110 | Near even, 6% apart |
| RosterAudit | 2,426 + 3,016 = 5,442 | 6,110 | Hampton side wins, 11% gap |
| Stats Guy | 1,666 + 2,323 = 3,989 | 5,145 | Hampton side wins, 22% gap |
On KeepTradeCut's scale, Jacobs plus Williams outsums Omarion Hampton by 40%. Its adjustment claws back 2,505 points for the Hampton side and lands the trade dead even. On FantasyCalc's scale the pair also outsums Hampton, and its fixed waiver credit closes most of the remaining gap. Both tools end within a few percent of even on a trade the community called a fleece. The two curve-based systems price the two aging veterans far enough below the young RB1 to name the winner, ours by a wider margin than RosterAudit's.
Trade 2, one-QB league, August 2026: Chase Brown, J.K. Dobbins, and Braelon Allen for De'Von Achane.
The replies: “Achane and it isn't really close.” Several replies were simply the one word “Achane.”
| Calculator | Brown + Dobbins + Allen | Achane | Verdict |
|---|---|---|---|
| KeepTradeCut | 6,163 + 3,724 + 3,345 = 13,232 | 7,345 + 2,284 adj = 9,629 | Brown side wins, 16% apart |
| FantasyCalc | 5,085 + 1,648 + 1,372 = 8,105 | 6,760 + 1,506 waiver = 8,266 | Near even, 2% apart |
| RosterAudit | 3,865 + 484 + 221 = 4,570 | 5,830 | Achane side wins, 22% gap |
| Stats Guy | 3,048 + 449 + 205 = 3,702 | 4,726 | Achane side wins, 22% gap |
Same shape, more extreme. On KeepTradeCut's scale the three-back pile outsums Achane by 80%, its adjustment recovers 2,284 points, and it still names the losing side the winner on a trade the community said “isn't really close.” FantasyCalc's waiver credit pulls the totals to within 2% and it calls the trade even. The aggregate numbers below are mostly this pattern repeating.
The results
RosterAudit is not in this three-way table. Its published values begin in June 2026, so it is scored on the window they cover, in its own section below.
Of the 3,004 trades priced by all three full-window systems, 2,803 ended with a clear or leaning community winner. Right when it names a winner measures how often the calculator's named winner matches the community's. Right answers delivered measures how often, across all 2,803 trades with a community winner, the calculator handed you that answer, where calling a decided trade even counts as a miss. Measured August 2026.
| Calculator | Right when it names a winner | Right answers delivered |
|---|---|---|
| Stats Guy | 85% on 2,058 named winners | 63% |
| FantasyCalc | 82% on 1,501 named winners | 44% |
| KeepTradeCut | 57% on 2,056 named winners | 42% |
When FantasyCalc names a winner it is nearly as accurate as we are, but it declines to answer on far more of the trades the community decided, so the three-point gap in the first column becomes a 19-point gap in the second. KeepTradeCut is the opposite. It speaks as often as we do and lands on the community's side just over half the time. Stats Guy is the only calculator in the study that answers “who won this trade” correctly more often than it calls the trade even or misses.
Fairness bands affect these numbers, so here is every calculator at every band from 4% to 16%, where the band is the gap, as a share of the larger side, inside which the calculator says even. A tighter band means naming a winner more often at some cost in accuracy, so each column is a dial between the two.
| Even band | Stats Guy, right / delivered | FantasyCalc, right / delivered | KeepTradeCut, right / delivered |
|---|---|---|---|
| 4% | 81% / 73% | 77% / 61% | 56% / 50% |
| 6% | 82% / 69% | 78% / 56% | 57% / 47% |
| 8% | 84% / 66% | 80% / 50% | 57% / 44% |
| 10% | 85% / 63% | 82% / 44% | 57% / 41% |
| 12% | 86% / 59% | 84% / 38% | 58% / 39% |
| 14% | 88% / 55% | 85% / 33% | 58% / 37% |
| 16% | 89% / 51% | 87% / 28% | 59% / 35% |
At every band, Stats Guy delivers more of the community's answers at equal or better accuracy than FantasyCalc, and KeepTradeCut's accuracy barely moves with the band at all, which says its gap comes from the values rather than the fairness rule. There is no setting of the dial at which Stats Guy is not first on both measures. The endpoint of the dial is no band at all, where every tool must pick a side on every trade, however thin its margin. Scored that way, Stats Guy picks the community's side on 79% of the 2,803 decided trades, FantasyCalc 73%, and KeepTradeCut 55%.
Split by trade type, the differences concentrate in the packages, where most dynasty trades are made.
| Trade type | Trades | Stats Guy | KeepTradeCut | FantasyCalc |
|---|---|---|---|---|
| 1-for-1 | 416 | 82% on 293 named winners | 78% on 313 named winners | 80% on 250 named winners |
| Several-for-one | 1,431 | 86% on 1,041 named winners | 46% on 1,062 named winners | 80% on 745 named winners |
| Multi-player both sides | 1,157 | 86% on 724 named winners | 63% on 681 named winners | 86% on 506 named winners |
On several-for-one trades, the most common shape in the sample, KeepTradeCut names a winner as often as we do and lands on the community's side less than half the time. On clean one-for-one swaps all three calculators perform respectably, and KeepTradeCut is the most willing of the three to make a call there. If your trades are strictly one player for one player, the tools are closer than the headline numbers suggest.
Two more checks. On the 556 trades where our calculator and KeepTradeCut's both named a winner and disagreed, the community sided with Stats Guy 469 times to 87. Open disagreements with FantasyCalc are rarer, 80 in the whole sample, and the community sided with us on 59 of them. And on the ~660 trades from August 2026, the newest month in the sample, the picture is the same, 84% on named winners and 63% delivered for Stats Guy, against 80% and 43% for FantasyCalc and 59% and 43% for KeepTradeCut.
Adding RosterAudit
RosterAudit launched in spring 2026, so its published value history begins May 31, 2026, and it is scored on the window its values cover. From June onward, 1,281 trades could be fully priced by all four systems, 1,194 of them with a community winner. Its calculator and ours use the same fairness rule, even within 10% of the larger side, so this is a direct comparison on the same trades, the same dates, and the same rule.
| Calculator | Right when it names a winner | Right answers delivered |
|---|---|---|
| Stats Guy | 85% on 898 named winners | 64% |
| RosterAudit | 80% on 826 named winners | 55% |
| FantasyCalc | 80% on 640 named winners | 43% |
| KeepTradeCut | 58% on 859 named winners | 41% |
RosterAudit is the second-best calculator in the study. It delivers more of the community's answers than FantasyCalc or KeepTradeCut, with real accuracy, and it does so with the same architecture we use, values learned from real trades and the package premium carried in the curve. Against us, at the same rule and near-equal willingness to make a call, Stats Guy is right more often, 85% to 80%, and delivers more of the community's answers, 64% to 55%. On the 677 trades where both calculators named a winner, we disagreed 49 times, and the community sided with Stats Guy on 35 of the 49. And with no fairness band at all, every system forced to pick a side on all 1,194 decided trades, the order holds: Stats Guy 79%, RosterAudit 74%, FantasyCalc 71%, KeepTradeCut 56%.
Two caveats in RosterAudit's favor that we cannot resolve from the outside. Its API serves daily value history for players but not for rookie picks, so trades containing picks, about seven in ten of this window, were priced with its current pick values rather than the values it showed on the trade's date. And unlike the other systems here, whose values we recorded ourselves throughout the window, we cannot independently verify that its served history matches what its site displayed on those dates. We did verify that its live calculator today reproduces exactly from its published values, sums, and band. Player values used in the RosterAudit comparison are by RosterAudit.com.
Tracking the strength of consensus, not just the direction
A good calculator should not just pick the right side. Its margin should be larger when the community is emphatic and smaller when the room is split. Because the community verdicts are built from counted voices, we can measure that directly, as the correlation between the calculator's margin and how lopsided the community's voices were. The correlation is 0.61 for Stats Guy, 0.52 for RosterAudit on its window, 0.47 for FantasyCalc, and 0.11 for KeepTradeCut, whose adjusted margins carry almost no information about how strongly the community will feel.
Four ways to build a calculator
Every calculator has to put the package premium somewhere. KeepTradeCut's calculator applies an adjustment at verdict time that favors the more consolidated side of a trade. FantasyCalc's credits the side receiving fewer players with a fixed waiver value. RosterAudit builds the premium into the values themselves, through an Elo-style rating over real trades, and its calculator compares plain sums. Ours works the same way, with the premium carried in the value curve, so our verdict is a straight comparison with no calculator-side correction.
Because we ran every system's real logic, we could grade each system's values with the calculator layer stripped away, as plain sums. The gap between the two numbers is what each tool's correction layer contributes. KeepTradeCut's raw sums are right on 43% of named winners and its adjustment lifts that to 57%. FantasyCalc's sums without the waiver credit are right on 61%, and the credit lifts that to 82%, twenty-one points of accuracy from a correction applied at verdict time. Our number is 85% either way, because the curve is the layer. Keep this in mind when judging league power rankings from different sites. Player trade value may be overstated in certain contexts.
Dynasty managers cite KeepTradeCut values constantly, because they are a very strong market signal. Yet when its calculator renders a verdict on the most common package shape, several-for-one, the same community overrides it more often than not. People trust the crowd's player values and then apply their own package math on top. That gap between the values people quote and the verdicts people reach is exactly what a trade calculator is supposed to close.
Where the other tools are strong
FantasyCalc's accuracy when it names a winner is close to ours, and if you prefer a tool that only speaks when it is confident, that profile is a reasonable choice. KeepTradeCut delivers more of the community's answers on one-for-one swaps than any other tool in the nine-month comparison, and its crowdsourced values respond to breaking news faster than any trade-data system can. RosterAudit's results back up the attention it has been getting. If you use any of these tools, you're using a good one. The differences show up in the packages, where most dynasty trades are made.
What we still get wrong
We are wrong on roughly one of every seven trades where we name a winner, and the misses are not random. When we miss, we lean toward the side receiving the single best player somewhat more often than the community does. One-for-one swaps are our least differentiated shape, and KeepTradeCut delivers slightly more right answers there than we do. The rest of the misses concentrate in cases that are hard for every value system. Verdicts driven by breaking news the trade market hadn't priced yet, verdicts about roster context like whether a rebuilder should be buying a veteran at all, and future picks where the community knows whose pick is being traded and prices in the roster it comes from. We price early, mid, and late versions of a future pick when the trade specifies one, but when a thread just says a 2027 1st and the replies all know it belongs to a rebuilding team, we miss out on important context. On those trades the other calculators rarely did better, and they set the current ceiling of the test.
Dated results, standing test
All four calculators keep evolving, ours included, so these numbers carry their measurement date: August 2026. A smaller run of the same comparison on about 700 trades in July 2026 reached the same conclusions, and the larger study confirmed them. New months of community trade debates arrive continuously and become fresh test data that nothing was tuned on. We'll keep running the comparison, and if the picture changes, this page will change with it.
Related reading: We Grade Our Verdicts Against Real Dynasty Communities and Packages, 1-for-1s, and the Value Curve.
