Investigate / Football / Outliers
How Unusual Was the Patriots’ Kicking Gap?
From 2001 through 2019, New England made 9.4 percentage points more of its field goals than its opponents. The next closest team didn't even get to four.
Result
After adjusting for distance, game situation, weather, and who was kicking, holding, and snapping, the gap shrinks from 9.4 points to 4.7. The remaining gap still leads the league. Opposing kickers tended to perform worse at Gillette than they did in other stadiums, but the data there are a little too underpowered to feel super confident.
- +9.4
- the gap in the original table.
- +4.7
- what's left after the full adjustment.
- -3.2
- how those same kickers did at Gillette versus elsewhere.
The league distribution
Holy crap are the Patriots an outlier here
Let's start with the original table. Pats kickers went 533 out of 618 (86.25%) on regular-season field-goal attempts. Opponents went 398 out of 518 (76.83%). This gives us a 9.41-point difference. Indy was second place at 3.81. The Pats' lead over the Colts was bigger than the Colts' lead over their own opponents. The Pats are 4.87 standard deviations above the average of the other 31 teams. I know what you're thinking. “Alex, that's a lot of standard deviations!” Yes, yes it is.
These are percentage points. Made, missed, and blocked field goals all count as attempts. I combined franchises that moved. The vertical line is zero.
Formal outlier check
The Patriots row looks like it wandered off from the rest of the chart. I ran a formal outlier test on it. A two-sided Grubbs test, to be exact. It returned p = 0.0013. For the pickers of nits: I picked the test after I had already seen the table, and the test assumes team results are roughly normally distributed. This was just me adding more quantification to our question.
Repetition
Maybe they took a few seasons off so it wouldn't be so obvious?
Nope. New England came out in the black in 14 of the 19 seasons. They were in the red in four, and dead even in one. They got a little carried away in 2002 with a ridiculous +31.2 points. 2019 was their worst season at −10.1.
I also ran everything again treating each season as one observation instead of each kick, since kicks from the same game and the same season aren't really independent of each other. The average season came out at +8.00 points with a 95% interval of 3.33 to 12.67 and p = 0.002.
One dot per regular season. The line connects the years so they're easier to follow. It isn't a fitted trend.
Expected make probability
Well, turns out who's kicking kinda matters
Not all field goals are created equal. A 60-yarder in Green Bay in December is not the same kick as a 25-yarder in Jerry World. I created a relatively simple model to help predict the likelihood of a field goal being made. The model knows the distance, season, home or road, roof, surface, temperature, wind, precipitation, quarter, score differential, time remaining, whether the game was close late, and how many kicks the kicker had already tried that day.
Then I added the kicker, holder, and snapper as their own effects. Kickers with a handful of attempts are weighted appropriately according to volume. Each game was also held out of the data used to score it, which means the model never got to peek at the first quarter before grading the fourth (data leakage to my fellow data scientists).
Each step adds another set of adjustments. The last row also removes blocked kicks, so it isn't another clean step down from the row above.
Full personnel model
Pats kickers came out slightly better than the model predicted at 0.16 points above prediction. Their opponents came in 4.58 points below prediction. After adding these together, 4.75 points of the original gap are still staring us in the face. This still leads the league over the observed period and is 3.42 standard deviations above the other teams (yes, still a lot of standard deviations). We did a little resampling of Patriots games, which put the range at 0.74 to 8.94 points.
I shrank the player effects more and less than the version I ultimately went with. The residual moved between 3.99 and 5.37 points. I then excluded all the blocked kicks from the data and it ended up at 4.09.
Out-of-game validation
The math seems to be mathing (I actually hate that phrase, I'm sorry)
To gain some confidence in what the model said about the Patriots, I looked into how it priced an ordinary NFL kick. I looked at all 18,461 attempts, and the model predicted an 82.46% make rate. The actual make rate was 82.36%. The out-of-game Brier score was 0.1245 and log loss was 0.3939. These measure how far the predicted probabilities are from the actuals. The lower the better. Area under the curve (AUC) was 0.771. This checks if the model puts easier kicks ahead of harder kicks.
If our model was perfectly calibrated, we'd have a slope of 1.00. Our model came in at 0.89, which means it was a little overconfident at the extremes. I think what we have here is fine. I just saw a tweet I thought was interesting and had some free time in the waning days of my paternity leave. Don't take this to the sportsbooks.
Each dot contains about 1,846 attempts. Dots on the diagonal performed exactly as the model expected.
The stadium hypothesis
Something in the water in (the visitor's locker room) in Foxborough?
Quality of opponent is figured into the raw home and away numbers. It changes week to week. The better comparison is looking at one kicker against himself. From 2012 through 2019, I recorded 35 opposing kickers who attempted a field goal at Gillette and another stadium. Together, they accounted for 117 non-blocked attempts at Gillette.
After we adjust for the kick itself, those 35 kickers were 3.18 percentage points worse at Gillette. I ran a game-level bootstrap, which ranged from −9.46 to +2.91 points with p = 0.319. Honestly, this doesn't really prove a “Gillette Penalty.”
-3.18 points95% interval -9.46 to +2.91
I removed blocks here. The Gillette attempts were 0.79 yards longer than these kickers' attempts elsewhere, on average. Distance and the rest of the recorded game context are included in the adjustment.
The broader home-road split
From 2001 through 2019, the raw gap was +10.02 points at home versus +8.87 points away. Looks like the Patriots' advantage traveled. The 2001 home games were at Foxboro Stadium, not Gillette.
Both bars use a zero-to-12-percentage-point scale.
What remains unobserved
Some limitations of this analysis...
Ryan Paganetti reported finding an offset image of the uprights on the Gillette video board during opponent attempts in broadcasts from roughly 2012 through 2020. As far as I know, no one is tracking this in play-by-play data, and I'll need a whole new paternity leave to focus on recording that data. All I have is where a kick happened. If someone wants to go through all the film, hit me up and I'll send you my code.
What the model can establish
A little p-hacking won't make it go away
As any good data scientist would, I started putting all sorts of thumbs on the scale to see what would hold and what wouldn't. After adjusting for everything I could think of (which is certainly distinct from adjusting for everything), the original 9.41-point gap drops to 4.75 and keeps the Pats safely in the lead. The Patriots' kickers did what the model expected. The remaining gap seems to come from opponent kickers performing below expectations.
Measurement and model details
I pulled every regular-season field goal in nflverse from 2001 through 2019. A “made” result counts as a make. “Missed” and “blocked” count as failures. I left out extra points and the playoffs. There are five weird plays marked as no-play because a penalty was enforced between downs. nflverse still records what happened on the kick, so I kept them.
Teams that moved stay together. That means San Diego and the Los Angeles Chargers are one franchise, St. Louis and the Los Angeles Rams are one, and Oakland and Las Vegas are one. For each franchise, I pooled all 19 seasons and subtracted its opponents' make rate from its own.
Distance isn't linear, so I modeled it with a six-knot cubic spline. The other numeric inputs are temperature, wind, score margin, seconds left, and the kicker's attempt number that day. I also included season, home or road, roof, surface, quarter, precipitation, and whether the game was close in the final five minutes. Kicker, holder, and snapper effects are shrunk toward league average. Holder and snapper names come from the written play description.
Kicker identity is available for every attempt. Holder identity is available for 99.99% and snapper identity for 99.93%. Temperature and wind show up 75.00% of the time. When they're missing, the model fills them in and keeps a separate flag saying they were missing. I tried ridge penalties of C = 0.1, 1, and 10. C = 1 had the best out-of-game log loss.
The validation split keeps complete games together. No game is partly in training and partly in testing. For the 4.75-point range, I resampled Patriots games while keeping the predicted probabilities fixed. I did the same for the kicker comparison, weighting each player by his Gillette attempts. Those ranges cover variation from the games in the sample. They don't capture every possible way the model itself could be wrong.
The analysis contract is patriots_field_goal_outlier_v2_2026_09_02. The Python code, frozen configuration, discovery notes, confirmation report, and publication review all live in the site repository.
Sources
The attempt data come from nflverse play-by-play releases and nflfastR field descriptions. nflverse publishes the data under CC BY 4.0. I checked New England’s 533-for-618 total against StatMuse. The original table appeared in this tweet. The video-board report is from Ryan Paganetti’s September 2, 2026 report.
Analysis and calculations: Alex Prejean. Data accessed September 2, 2026.
Conclusion
Okay, so if you don't like numbers and data and skipped straight here (why are you even reading this?), we started with the 9.41-point figure, adjusted for factors related to the kick, and ended up with a 4.75-point difference, which is still a big number. If we drop blocked field goals, it drops further to 4.09, which is still a big number. We can statistically say it's a big and unexpected number. The only thing we can't say for sure is what sort of Patriot witchcraft caused it.