Lately I've been building a fantasy football decision engine, one model at a time. One of the newest pieces is trade evaluation. Most trade models compare market value to market value, maybe with a discount when one side sends more pieces than the other. Almost none of them are specific to your league.
A lot goes into a dynasty trade:
- Outlook: are you rebuilding, contending, or somewhere in the middle?
- Volatility: a player coming off a breakout season may not sustain it.
- Long shots: the upside bets that rarely hit, but change a roster when they do.
- Picks: future value with no player attached yet.
Common trade calculators and dynasty models mostly ignore this calculus. That leaves it to you to either understand the system perfectly or trust your gut about where your team is headed and what to prioritize.
There's a reason most calculators skip these factors. A market value that's the same for every team doesn't have to answer for nearly as much error or bias. Once a model says a trade is good for your team, it has to be right about your team. So the questions become: how do we find and reduce the error, and how do we show people they should trust the model with their roster?
The model I've built values a trade by the change in your team's playoff odds over the next five seasons. I grade the first three, the ones I can check honestly. When I look for error, I'm looking for where the model led you astray: not where a trade returned less than you expected, but where it was clearly wrong and made your team worse. You can live with a trade that doesn't return as much as you hoped. Taking one that cripples your team is a real loss. And any fair evaluation has to allow for what no model sees coming: injuries, other trades, breakouts, and situations that change.
The unit. Everything here is measured in points of playoff odds. A trade worth +10 raises your chance of making the playoffs by 10 percentage points, summed over three seasons.
Building the dataset
For the test bed I generated 21,600 random trades in simulated leagues built from the 2021–23 drafts. That's 43,200 trade sides, one for each team in a trade. Each trade was replayed against the real seasons that followed: 16 replays of 1,024 simulated seasons each, so the luck of any single replay averages out. I'm keeping the generation details light on purpose, since there are a lot of ways to build these. The mix covers lopsided trades, "fair" trades, picks-only deals, 1-for-1 swaps and multi-player packages. Because the trades are random, the model never chose any of them.
The test bed
43,200 trade sides, one per team in each trade
- Simulated leagues2021–23 drafts
- Random trades21,600 trades
- Model predictsΔ playoff odds
- Replay vs. real seasons16 × 1,024 sims
- Realized resultper trade side
- Error scorepredicted vs. realized
Two rules kept the search honest:
- Find it in one year, confirm it in the other two. Each draft year takes a turn as the discovery year. A pattern found there is frozen, thresholds and all, and has to show up again in both of the other years.
- Count players, not rows. Thousands of trades can involve the same 20 star players. Every pattern also has to survive resampling by player. One that only survives resampling by league gets labelled few-player: promising, but carried by a handful of careers.
Discover once, confirm twice
| 2021 draft | 2022 draft | 2023 draft | |
|---|---|---|---|
| Round 1 | Discover & freeze | Must confirm | Must confirm |
| Round 2 | Must confirm | Discover & freeze | Must confirm |
| Round 3 | Must confirm | Must confirm | Discover & freeze |
One more number sets expectations. The model misses a typical trade by about ±26 points, while the replay's own noise is about 4. Most of that miss is how careers actually went, which nobody could have known at the time of the trade. So most real patterns in this piece are small, a few points either way. The handful of big ones turn out to be carried by a few star careers.
In my research I came across four interesting ways to find error patterns in a set of results. The idea started with applying k-means clustering to the error surface, to find pockets of error and decide what to work on next. From there I found the shallow decision trees behind Microsoft's Error Analysis tool, and two slicing methods, SliceLine and SliceFinder. All four use some mix of trees, clustering or combinatorics to find what's actually causing the errors. Below I'll go through each one, what it found in the model, and how I'd use it in the future.
K-means clustering
K-means works like this:
- Take just the error rows.
- Standardize all of their facts (ages, scoring levels, market prices, team strength, and so on).
- Let k-means find 3 to 8 natural clusters, keeping whichever count separates them best.
- Put every trade, not just the errors, into its nearest cluster.
- Ask: does any cluster have a higher error rate than the rest?
If the error rate is 15% across the board, does any cluster run at 20%, 30% or 40%? That would suggest something about the cluster is a real contributor to the error.
I ran it for every error target and year, about 90 clusters in all, and there was no real cluster structure. Separation scores sat around 0.1, where 1 means crisp clusters and 0 means none. A couple of clusters did hold up out of year. The clearest was trades with high prices on both sides, which ran at about 1.5–1.7× the normal error rate. But the slicing methods found the same group with a simple one-line rule, which makes k-means mostly redundant here.
Error enrichment of the three k-means clusters, found on 2023
High prices on both sides · held every year
Low scorers sent, weaker team · few-player
Strong team, weaker partner · failed
Part of the problem is how k-means works. It tends toward compact, similar-sized clusters, and the clusters never see the error while they're being built. It isn't wrong to group trades that way; it's just not grouping them by what went wrong, which is what the other three methods do.
Shallow decision trees
A shallow decision tree, the engine of Microsoft's Error Analysis tool, takes the whole dataset and picks the single fact that best splits the rows into a high-error half and a low-error half. Then it does the same inside each half, three levels deep.
In the tree below, the first question was is the best player we sent averaging 21 or fewer points per game over the last three years? For the trades where that's yes, the next question was did he score 17.8 or fewer last season? Each answer narrows the trades down until every leaf at the bottom has a clearly higher or lower error than average.
The actual tree, discovered on 2023
Each leaf's average miss (realized − predicted, playoff points) in 2023, then in the two years the tree never saw.
Run on bad takes instead, a tree found that trades sending away a non-QB who scored more than 15.4 points per game last season went bad 56% of the time in 2021 against a normal rate of about 30%. It stayed high in the two years the tree never saw: 46% in 2022 and 59% in 2023. That's a pattern worth pulling against. It was also few-player, carried by about 25–34 players a year, so it's a strong lead rather than a proven fix.
That out-of-year check is the whole point. If a leaf isn't also elevated in the years it wasn't fitted on, I can't tell a real error from something I got lucky finding in the discovery year. The 17.8+ ppg leaf above is exactly that case: +10.0 where it was found, negative in both other years.
The weakness of the tree is that it's greedy. Once it commits to a first split, a real error group that straddles both the yes and no branches gets cut in half and may never surface. That's what the slicing methods are for: they test every one- and two-condition combination, so they don't miss a question just because it didn't come first.
SliceLine
SliceLine cuts every fact into quartiles, or keeps small counts as they are ("we sent 0 quarterbacks", "we sent 1"). Every single condition and every pair of conditions becomes a slice, and every slice gets a score: how much worse its average error is than everyone else's, minus a penalty for being small. The balance between "how much worse" and "how big" is a setting called alpha, and there's a minimum of 300 trades per slice. Without that balance, a group of two with a 100% error rate would look as important as a group of a hundred running at 20% against a 10% baseline.
That came to about 42,000 slices per error target and year. SliceLine only works with errors that are never negative, so the model's miss was split into two searches: an over-statement half (the trade came in worse than predicted) and an under-statement half (it came in better). The top 10 slices from each search were carried to the other two years.
With 2023 as the discovery year, the #1 over-statement slice was trades where we send no quarterback and the best player we send scored more than 16 points per game last season. It held in every year:
SliceLine's #1 slice: sends no QB, best player sent scored 16+ ppg
2021 · model off by −9.1
2022 · model off by −7.7
2023 (found) · model off by −5.5
SliceFinder
SliceFinder builds the same kind of slices, but treats each one as a hypothesis test. It compares each slice to the rest with a t-test, and the slice has to clear an effect-size floor (Cohen's d of at least 0.2, "small but real"). Thousands of tests will throw up false alarms, so the false-discovery rate is held at 5%. It starts with one-condition slices and only combines the ones that weren't already flagged, so the simplest slices come first.
In 2021, SliceFinder examined 29,486 slices, and 859 passed. The two largest were telling:
- We receive players whose comparable players project 10–15 points per game in total. A big miss in 2021, barely there in 2022, and the opposite in 2023.
- We send no first-round pick, and the best player we send scored more than 17 points per game. A clear miss in 2021 and a bigger one in 2022, then almost nothing in 2023. No signal.
Tiny p-values, no repeat
Receive players projected 10–15 ppg total
Send no 1st, best player sent 17+ ppg
SliceFinder didn't confirm a single bias slice. The lesson is sharper than "it's strict": a tiny p-value in one year is not evidence that a pattern repeats. Each season has its own character, and a test inside one season can't see that. Where an effect really is broad, though, SliceFinder was the most reliable of the four. Every one of its 30 slices on the size of the miss confirmed, nearly all of them versions of "big trades miss by more".
What did we find?
Combining all four methods, the biggest lesson was that big trades miss by more. No surprise there: the more value moving, the more room for something to go wrong. Setting k-means aside, the tree, SliceLine and SliceFinder each found a version of the same bias: the model under-values elite non-QBs. It expects them to fade over the next two seasons, and the best ones mostly didn't. It's essentially one bias, carried by about 20–45 top-tier players.
That felt like an answer. It wasn't the right one.
The wrong question
Every method had been working off the same definition of error: the gap between predicted and realized value. With trades, error isn't that black and white. If the model says a trade is worth +16 and it delivers +9, winning 9 is still valuable. In the context of the trade, it can be worth more than holding on to what you had, and more than not making the trade at all. Going from +3 to −5, or even to −1, is the real error, the one I actually want to find.
Here's a trade from the replay. We get Stefon Diggs, Phillip Lindsay and a first; we give DJ Chark, Rhamondre Stevenson and a first. The model said +17.7. It came in at +10.0. By the measure every method had been using, that's a 7.7-point error.
Here's another. We get Ryan Fitzpatrick plus picks for Kareem Hunt. The model said +2.7. It came in at −3.2. That's a 5.9-point error, smaller than the Diggs trade.
Same two trades, two measures
Diggs + Lindsay + 1st · model +17.7 → real +10.0
Fitzpatrick + picks · model +2.7 → real −3.2
But the Diggs trade was still a good trade. We'd make it again. The Fitzpatrick trade is the one that led us astray. Once I went looking, 57% of everything I'd been calling "error" came from trades where the model made the right call.
Where the "error" was hiding
Share of each measure, by the model's predicted value (playoff points over three seasons)
| Model's call | Share of all miss | Share of all decision cost |
|---|---|---|
| Below −10 | 33.0% | 18.4% |
| −10 to −5 | 7.9% | 11.0% |
| −5 to −2 | 5.2% | 10.1% |
| −2 to −0.5 | 2.4% | 5.7% |
| Within ±0.5 | 1.7% | 4.4% |
| +0.5 to +2 | 2.6% | 6.1% |
| +2 to +5 | 5.2% | 11.1% |
| +5 to +10 | 8.3% | 12.0% |
| Above +10 | 33.8% | 21.2% |
So I changed the question. Not "how far off was the number?" but two new ones:
What did the call cost?
How much better the other choice would have been, where not trading is worth 0. A +16 that delivers +11 costs 0. A +2 that lands at −2 costs 2. A +5 that lands at −5 costs 5.
Was it wrong more often than it should be?
How often a call lands on the wrong side of zero, minus how often a call that size normally does. A +0.5 that flips is a coin flip and scores about zero. A +5 that flips scores high.
The question stopped being "how far off was the number?" and became "what did the decision cost?"
So what did we change?
We changed how we score a mistake. The model itself hasn't changed yet. Every method re-ran on the two new scores, and every pattern from the first pass was re-tested. The order changed:
Before and after
To make each finding concrete, here are trades from the replay. They're anecdotes, picked because they went wrong; the rates beside them are the evidence.
1. Your team's read vs. the market
The model values a trade for your roster. The market values it for an average team. When the model says take it but the market says you're overpaying, the take goes clearly bad 31% of the time, against 15% overall, and costs more than twice the average. This was the #1 decision error in every year, under both resamples.
Share of "take" calls that ended clearly worse than not trading
2021
2022
2023
2021: we get Dak Prescott, Darrell Henderson and a pick; we give Marquise Brown, Austin Ekeler and a first. Ekeler broke out.model +8.5 · market −89 · realized −30.0
2. Depth, and mismatched trades
Getting two or three players for one is over-credited, mostly in the season of the trade. The model counts the extra bodies as if they'd start. This held in every year under both resamples.
2021: we get Kellen Mond, Michael Carter and Elijah Mitchell, all rookies; we give Mark Andrews.model +31.6 · realized −27.4
2021: we get JuJu Smith-Schuster, Darnell Mooney and a pick; we give Mike Evans.model +26.5 · realized −27.4
3. Rookies
Receiving a rookie costs more than average in every year. The first pass missed this entirely. Rookie misses go both ways, so on the old measure they averaged out to nothing. As decisions, both directions cost.
2022: we get Jameson Williams, Baker Mayfield and a pick; we give Saquon Barkley and two picks.model +13.1 · realized −26.4
4. Big trades: the same call means less
The first pass said big trades miss by more. The second pass says something more useful: the model is over-confident on them. Hold the size of the model's call fixed and only change how much is moving, and the wrong-side rate roughly doubles:
How often the call lands on the wrong side
By the size of the model's call (either sign) and how much value is moving, in fifths of all trades
| Model's call | Smallest fifth | 2nd | 3rd | 4th | Biggest fifth |
|---|---|---|---|---|---|
| 0 to 2 points | 18% | 26% | 25% | 33% | 38% |
| 2 to 5 | 18% | 22% | 26% | 30% | 37% |
| 5 to 10 | 14% | 17% | 21% | 21% | 27% |
| 10 to 20 | 11% | 12% | 13% | 17% | 21% |
2021: the model said decline. We'd have gotten Devin Singletary and Mark Andrews, and given Russell Wilson, Ezekiel Elliott and a pick. Andrews broke out; Wilson and Elliott declined.model −37.8 · realized +34.3
The fix isn't a bigger or smaller number. It's accepting a trade on the chance it gains, which needs ranges that widen with the stakes.
5. The elite-player bias, re-ranked
The first pass's headline is still real in points. The model expects elite non-QBs to fade in years two and three, and many don't. But it rarely flips a decision (15.8% wrong side against 15.1%), because those calls were usually clear anyway. Where it matters is how much to ask for.
2021: we trade Davante Adams for Jamison Crowder and two firsts
2021 (trade season)
2022
2023
What the model gets right
Picks-only deals land on the wrong side 8.3% of the time against 15.1% overall, at less than half the average cost. Lopsided trades, both the clear wins and the clear losses, cost less than average too. The model is most reliable when the picture is simple.
One wrinkle. Searched on its own, raw decision cost mostly rediscovers which calls are hard: trades that are close by market value, and trades where a lot is moving. That's the difficulty of the task, not the model's mistake. Surprise corrects for how big the call was, so it became the score to search on.
What is each tool good for?
So when should you use each of these? Here's what each one turned out to be good at:
Shallow decision tree
The most readable first look. Every leaf is a rule you can say out loud. Watch for its greed: a real group split across branches can disappear.
SliceLine
Best at two-condition groups, like "big on both sides of the trade". Its top 10 are often one group described ten ways, so merge before you count.
SliceFinder
The most reliable for broad effects, but it stops at one condition whenever one condition is enough. It found where the model is safe, not where it errs.
K-means
Weak here. It groups trades by resemblance, never by error, so it's solving a different problem from the other three.
Your own hypotheses
Write the question down first, then test it. It's slower, but the two strongest findings here (team read vs. market, and depth) came from tests like this, not from a search.
Conclusions
The biggest thing I took away is to define the error before you go looking for it. Decide what counts as a mistake and how bad each one is, before choosing which samples to generate and which methods to run. Measure the decision, not the number. The first pass here wasn't wasted, but it was answering a question I didn't actually care about, and a clear definition up front would have saved a lot of digging.
Three smaller lessons came with it:
- Confirm out of sample. Patterns with tiny p-values in one season fell apart in the next. Discover once, confirm twice.
- Count players, not rows. Some of the biggest "errors" were really twenty careers. Resampling by player tells you which.
- Show the trades. A rate convinces no one until you can point at Mark Andrews for three rookies.