Saturday, August 1, 2026

Noah's Quest

HTML Genesis: The Building of the Ark

Noah: Faith & Obedience

Gather materials, board the animals, and build the Ark according to the pattern (Genesis 5-8)

Refining & Crafting
Ark Construction
Current Phase: Frame
Captured Animals
Clean Land (7 needed): 0
Unclean Land (2 needed): 0
Clean Birds (7 needed): 0
Unclean Birds (2 needed): 0
Journal

Saturday, July 25, 2026

Yahtzee Simulator - Beta

HTML Yahtzee - Custom Pip Debug Mode

YAHTZEE

Roll to start!
Category Score

Tuesday, September 3, 2024

Computer Vision

The delightful thing about Data Science in the novelty of it all. Classical Statistical Methods are time-tested, but not very sexy. For me, the thrill of discovering something new is to share it with others. As such, I thought it useful to write up a discovery during a recent project. The Aarhus University Signal Processing group, in collaboration with the University of Southern Denmark, has provided the data containing images of unique plants belonging to 12 different species. I was tasked to build a Convolutional Neural Network model which would classify the plant seedlings into their respective 12 categories:
I did some research, and the literature suggested that VGG19 does well with plant classification models. Using the adam optimizer and categorical cross-entropy loss the resulting model resulted in 60% accuracy in the training set, but only 45% accuracy in the test data, even after employing gaussian blurring to regularize the provided plant images. This was clearly not good enough.
Looking for a quick and inexpensive improvement, I turned to a project involving a CNN for the MNIST dataset in a hand-writing classification problem that was extremely discerning. My thought process was that plants, like handwritten characters, originate from a root/starting point and "grow" from there. Each plant is as unique as the digits 0 to 9, with minor variations due to environmental factors, much like the handwriting style of different individuals. The resulting CNN achieved 90% acccuracy in the training data and 70% data in the test set, after 30 epoch, with the apparent overfitting beginning to occur around 15 epochs
Digging in a bit to the test data, it was evident that the model has a challenging time discerning between imamature (unsprouted) Black-grass and Loose-Silky Bent grass, which is clear from the professional illustrations below (not from dataset we're analyzing).
Digging in a bit to the images themselves, it became clear that the grasses themselves were nearly indistinguisable prior to flowering, which helps explain why the model has a tough time with discernment, just as a human would:
In a similar vein, it's problematic how often the Loose Silky-Bent grass is confused with common wheat (at an approximate 50-50 rate, well, 17-17 technically). This suggests that further analysis would be impractical: it's better to let the proverbial "tares grow up among the wheat", just as Jesus preached.

Wednesday, August 14, 2024

Top 10 Posts

In updating my resume I had occasion to compile and organize some of my favorite sports analytics posts, both here and on external sites.  Here's the top 10, in chronological order:

Revisiting the Homecourt Advantage in College Basketball, the Sports Collection*, January 2012

Valuation of an NFL Pro-Bowl Tight End, Sportistician, June 2015

Effect of Size on NFL combine DrillsSportistician, March 2016

Short-Yardage Conversion Rates, Sportistician, March 2016

NFL Tight Ends and the Combine, Football Outsiders, March 2016

NFL Pro-Bowl Odds CalculatorSportistician App, June 2016

New Touchback Rules and Their Effect, Sportistician, June 2016

Why, Cy Young, Why?, Sportistician, Nov 2016

Tight End Prospecting, Rotoviz*, March 2017

2017 Tight End class, Rotoviz*, April 2017

*Indicates introduction/abstract available at no cost

Friday, July 7, 2023

NBA shot Clustering Analysis

Article courtesy of Avery Caraway

The game of basketball has changed dramatically over the last few decades. Teams are scoring at a much higher rate than in the past. In the 2005-2006 season, NBA teams scored on average 97.0 points per game [1]. In the 2020-2021 NBA season, the league has increased this total point average to 112.1 points per game [2]. One reason for this drastic growth is the increase in 3 point shots taken. In the 2007-2008 season, teams averaged 18.04 3-point attempts per game, but that number rose to 28.98 attempts per game in the 2017–2018 season [3]. Even if the field goal percentage remains the same, attempting more shots equals more points being scored, and in part, a better chance to win the game.

Many players in today’s NBA have modeled their game after the ability to make 3-point shots. Although teams are made up of more than just one individual player, the increasing rate of the 3 point shot has brought success to these teams. For example, let’s take a look at Stephen Curry of the Golden State Warriors. Now highly regarded as the best 3 point shooter of all time, he did not start off that way in the NBA. In his first year in the league, 2009, Curry attempted 4.8 3-point shots per game and was successful 43.7% of the time. In the year 2016, Curry led his team to 73 wins, the highest regular season total of all time. During this season, Curry attempted 11.2 3-point shots per game and was successful 45.4% of the time. Although his field goal percentage rate barely improved, the amount of shots taken drastically increased by over 6 attempts per game [4]. There are many characteristics of a great team, but one that can take advantage of the 3 point line has shown to be successful in recent years.   

To see how much shot locations have changed in the NBA, we can look at the Houston Rockets and San Antonio Spurs from the 2007 and 2017 seasons. In 2007, the Rockets had 7304 shot attempts and in 2017 it increased to 8698 attempts. Whereas, in 2007, the Spurs had 8075 shot attempts and in 2017 it increased to 8742 attempts. The game has increased in pace over the last 10 years with the Rockets in 2017 taking approximately 17 more shots a game and the Spurs taking approximately 8 more shots a game when compared to their 2007 season statistics


Does this change of pace and increased number of attempts have an impact on the location where the shots were taken? There are three different clustering techniques to analyze the shooting data; K-means clustering, gaussian mixture, and DB scan. These methods show differences in the shooting patterns of both teams.  

K-Means is a distance-based algorithm which clusters points based on closeness. This process requires the user to provide a number of groups to cluster the points in. The centroids of each of these clusters are generated randomly. These get adjusted every iteration to find the actual centroids of the clustering of data. It cannot however create irregularly shaped clusters.



Gaussian Mixtures are probabilistic models which use a soft clustering approach. It assumes that there is a certain number of Gaussian distributions. The clusters have a specific and unique mean and variance. The values of mean and variance are determined by ExpectationMaximization technique. This method does need the user to mention the number of clusters.


The DB scan method is a density based clustering algorithm that works on the assumption that clusters are dense regions in space separated by regions of lower density. It is very robust to outliers and can create irregularly shaped clusters. The user does not need to specify the number of clusters for this method. However, the circle radius from each data point and the number of points inside each circle from data points need to be provided. The number of clusters is determined through the iterations.


In both the Gaussian mixture and DB scan these models were able to create separate clusters of shots along the 3 point line. However, the k-means was unable to do so for the shooting data, because this method cannot cluster irregular shapes, as K-means calculates the distance from a centroid, which naturally forms spherical clusters. As such, K-means is a useful clustering method when researching players that have similar styles or roles. It is not as successful when clustering shots since it doesn’t create separate clusters around the 3-point line, like DB scan and Gaussian mixture do. Gaussian mixture was also able to show more shooting density in 2017 for both the Rockets and the Spurs under the 3 point line when compared to the 2007 season. Overall, the Rockets shot more 3 pointers and less mid range shots in 2017 when compared to the 2007 shooting data. On the other hand, the San Antonio Spurs had a similar shot distribution location in 2017 as it was in 2007.

Recommendations:
The next step in this research would be to see how much effect on overall winning percentage this increases in 3 point attempts has. It was mentioned above that the Golden State Warriors had the best statistical season of all time, largely due to their 3 point shooting ability. However, is this one team an anomaly, or will there be a true shift in the NBA game over the next 10 years? Should teams invest time and resources into putting 5 highly efficient shooters on the court, as opposed to the traditional Guard-Guard-Forward-Forward-Center lineup that fans are used to seeing? There are many factors involved in the overall winning percentage of a team, but it would be interesting to see how important the 3 point shot is compared to the other statistics. Doing an analysis on teams in the playoffs vs. teams not in the playoffs over the last 5 years may give information into this topic.

REFERENCES
1. “NBA league average points per game 2006”. StatMuse. 2006 https://www.statmuse.com/nba/ask/nba-league-average-points-per-game-2006
2. “NBA league average points per game 2021”. StatMuse. 2021 https://www.statmuse.com/nba/ask/nba-league-average-points-per-game-2021
3. Andrew Lisa (2019). “25 ways the NBA has changed in the last 50 years”. Stacker. https://stacker.com/basketball/25-ways-nba-has-changed-last-50-years
4. “Stephen Curry”. Basketball Reference. https://www.basketball-reference.com/players/c/curryst01.html

Sunday, March 1, 2020

2020 SAVVAGE SCORES

Using the logistic regression model for estimating the odds of a legitimate pro-bowl nomination as featured by Football Outsiders in 2016, we can use a linear model to estimate the log pro bowl odds of every tight end combine participant from the 2020 drills. With these players' pro-day data coming in at a later date, there's a good chance that a couple players clear the 6% success threshold (as established by historical data). For example, assuming Albert Okwuegbunam can manage at least a 28 inch vertical at his pro day, we can conclude he's likely to clear success threshold of this model.

Name Height Weight Vertical 40 time Vert+Hgt Weight/40 log odds P(AP1)
Cole Kmet 78 262 37 4.7 115 55.7 0.35 26.20%
Albert Okwuegbunam 77 258 28* 4.49 105 57.5 0.07 6.54%
Stephen Sullivan 77 248 36.5 4.66 113.5 53.2 0.06 5.95%
Adam Trautman 77 255 34.5 4.8 111.5 53.1 0.04 3.50%
Dalton Keene 76 253 34 4.71 110 53.7 0.03 3.28%
Colby Parkinson 79 252 32.5 4.77 111.5 52.8 0.03 3.00%
Dom WoodAnderson 76 261 35 4.92 111 53.0 0.03 2.97%
CJ O'Grady 76 253 34 4.81 110 52.6 0.02 1.84%
Brycen Hopkins 76 245 33.5 4.66 109.5 52.6 0.02 1.60%
Devin Asaisi 75 257 30.5 4.73 105.5 54.3 0.02 1.48%
Charlie Woerner 77 244 34.5 4.78 111.5 51.0 0.01 1.18%
Harrison Bryant 77 243 32.5 4.73 109.5 51.4 0.01 0.85%
Josiah Deguara 74 242 35.5 4.72 109.5 51.3 0.01 0.81%
Charlie Taumoepeau 74 240 36.5 4.75 110.5 50.5 0.01 0.70%
Hunter Bryant 74 248 32.5 4.74 106.5 52.3 0.01 0.66%
Mitchel Wilcox 75 247 31 4.88 106 50.6 0.00 0.24%
*hypothetical pro-day value

While this model is still in the validation phase, we've had some promising hits in players like Travis Kelce, George Kittle, and Tyler Eifert (out of 25 players clearing the threshold). The fact that this model has identified the *only* two Associated Press 1st Team Pro bowl Tight ends since the inception of the model suggest this model is very good at identifying high upside tight ends.

Though it's moving the bar somewhat, the associated ordinal test (AP 1st team, 2nd team, no pro bowl) comparing eventual pro bowls of combine participants above the 6% threshold to those below does yield a statistically significant p-value (though the assumptions of that test may be a bit questionable due to the small number of successes involved).

Applying this to my dynasty football league, I'll be stashing players like Kmet, Okwuegbunam, and Sullivan and looking to acquire predicted successes from last year's combine like T.J. Hockenson, Noah Fant, Foster Moreau, and Kahale Warring.

Wednesday, June 12, 2019

Scott Fish Bowl's Top 100 from 2018

The 2019 Scott Fish Bowl scoring was just released earlier this week. Assuming no changes, here is the scoring breakdown...

Passing:

    • 4 point passing TD
    • -3 point interception
    • -1 point interception for TD
    • 1 point for 25 yards passing (.04/per),
    • 1 point per 2 point conversions
    • .25 points per 1st down

Rushing:

    • 6 point rushing TD
    • 1 point for 10 yards rushing (.1/per),
    • 2 points for 2 point conversions
    • .5 point per 1st down

Receiving:

    • 6 point receiving TD
    • 1 point for 10 yards receiving (.1/per),
    • 2 points for 2 point conversions,
    • .5 point per 1st down
    • .5 point per reception

TE:

  • Extra .5 point per first down
  • Extra .5 point per reception

Returns:

    • 6 point for any return TD
    • 6 points if your player recovers a ball in the endzone for a TD
If memory serves, the scoring system is not drastically different than that of 2018, with the exception of the -3 point INT.

First down data is a bit tricky, but thanks to the play query tools at Pro Football Reference we can pull it fairly easily. The down side is that the lower probability events (like pick 6, fumble lost, and punt and kick return TDs) are a lot of trouble to query separately and merge with the player list. So if we neglect the splitting hairs of these low probability events, we can generate the hypothetical scores of the 2018 under this slightly modified scoring system.

Rank, Player, Score
1 Patrick Mahomes  476.58
2 Matt Ryan  421.71
3 Deshaun Watson  396.26
4 Ben Roethlisberger  388.89

5 Todd Gurley  387.60
6 Christian McCaffrey  382.75
7 Saquon Barkley  380.80

8 Andrew Luck  376.49
9 Aaron Rodgers  370.83
10 Jared Goff  365.32

11 Travis Kelce  364.60
12 Drew Brees  364.43
13 Alvin Kamara  352.90
14 Russell Wilson  347.85
15 Dak Prescott  344.34

16 Zach Ertz  343.80
17 Kirk Cousins  341.92
18 Ezekiel Elliott  341.60
19 Cam Newton  330.40
20 Tom Brady  330.31
21 DeAndre Hopkins  320.50
22 Tyreek Hill  320.30

23  Philip Rivers  319.52
24 George Kittle  318.20
25 Julio Jones  313.90
26 Davante Adams  304.60
27 Antonio Brown  303.70

28 Mitchell Trubisky  302.77
29 James Conner  294.50
30 Michael Thomas  294.50
31 Adam Thielen  290.30

32 Eli Manning  285.21
33 Melvin Gordon  281.00
34 Mike Evans  278.40
35 Baker Mayfield  275.89
36 JuJu Smith-Schuster  271.60
37 Eric Ebron  264.70
38 James White  264.60
39 Matthew Stafford  263.93
40 Robert Woods  261.10
41 Joe Mixon  260.40
42 Derek Carr  259.84
43 David Johnson  251.40
44 Kareem Hunt  250.70

45 Case Keenum  249.44
46 Keenan Allen  245.60
47 Jameis Winston  243.48
48 Josh Allen  243.11
49 Carson Wentz  241.60

50 Stefon Diggs  239.80
51 Jared Cook  238.10
52 Phillip Lindsay  235.50
53  Brandin Cooks  233.70
54 Tarik Cohen  230.59
55 Chris Carson  229.90

56 T.Y. Hilton  228.90
57 Derrick Henry  224.31
58 Odell Beckham  220.84
59 Blake Bortles  217.55
60 Marcus Mariota  216.02
61 Tyler Lockett  213.70
62 Nick Chubb  213.40
63 Adrian Peterson  213.00

64 Tyler Boyd  210.10

65 Tevin Coleman  204.60

66 Jordan Howard  203.50
67 Lamar Jackson  201.54
68 Kenny Golladay  201.10
69 Jarvis Landry  200.87

70 Andy Dalton  199.04
71 Sam Darnold  198.35

72 Marlon Mack  198.10
73 Calvin Ridley  197.00
74 Amari Cooper  196.50
75 Julian Edelman  194.67

76 Kenyan Drake  193.70

77 Austin Hooper  192.00
78 Aaron Jones  187.90
79 Ryan Fitzpatrick  187.34
80 Emmanuel Sanders  187.17
81 Lamar Miller  184.60
82 Kyle Rudolph  183.40
83 Chris Godwin  180.70
84 Larry Fitzgerald  180.13

85 Matt Breida  178.40
86 Trey Burton  177.60
87 Mike Williams  176.70
88 Alshon Jeffery  176.50
89 Corey Davis  175.60
90 Adam Humphries  175.20

91 Austin Ekeler  174.30
92 Mohamed Sanu  174.15
93 Sterling Shepard  172.00

94 David Njoku  171.90
95 Alex Smith  170.50
96 Ryan Tannehill  169.61
97 Joe Flacco  169.00

98 Sony Michel  168.10
99 Vance McDonald  168.00
100 Rob Gronkowski  167.60

Granted, this list neglects fumbles, which may be a larger factor for some players than others, it's still a decent list to see the trends of how the points scored would have broken down by position last year. With 14 QBs in the Top 25, it's evident that a late QB strategy would have paid dividends last year. Also, with only 3 tight ends in the top 25 and a large drop in total points to the 2nd tier tight ends, it's clear Kelce, Ertz, or Kittle won a lot of SFB8 leagues last year. Furthermore, a top 5 running back was a great help as well in 2018: Elliot, CMC, Barkley, Kamara, and Gurley bested their positional peers by a good 50 points or more. As for the Wide Receivers, there were nine WR1s within 50 points of each other, with the next tier of WRs containing about twelve WR target leaders for their team. In summary, a first round running back pick, a second round WR, third round TE likely would have been a very strong start in the Scott Fish Bowl last summer. Lucking out with some combination of top 7 QBs with later round picks, that hopefully included Patrick Mahomes, would have also paid dividends. Filling in the other rounds with surprisingly productive pass catching RBs and slot WRs could have easily led to a championship run.