The rundown
(in no particular order)
- 1.
The Youtube channel Business of Sport posted a video interview with Como's President, Mirwan Suwarso, back in July, but it's been doing the rounds on social media again this week. Rarely do you see an executive at that level provide such in-depth insight into the analytical processes happening behind the scenes, so it's an interesting listen — from Cesc's owner presentations to Billy Beane's influence (here's the relevant part).
- 2.
This week, a research paper titled "Earnings of Female Coaches in European Soccer" and based on private UEFA survey data was published in the Journal of Sport Economics. I found it somewhat disappointing in that the results lacked much applicability — the survey group included amateur coaches and their main significant finding was that weekly coaching hours correlated with earnings (yeah, I get paid hourly too) — but I do think that this kind of research is desperately needed, and hopefully this paper will lead to more in the future. (I also wish they made the data public, but alas)
- 3.
John Muller did an excellent job of breaking down Cavan Sullivan's unbelievable breakout season for The Guardian. The article is an excellent primer on Sullivan and his play style for those who aren't familiar; John writes "Sullivan’s style on the ball is less a young Messi and more like his other childhood hero, Alexis Sánchez: hips low, elbows out, juking defenders with feints, stutters, stepovers and changes of direction." He's also one of the league's best progressive passing attackers, by the way, and has 0.48 xG+xA per 90.
- 4.
MLS is hiring a Senior Manager, Soccer Analytics and Data Science. There's some business stuff in the job description but also things like "lead the evaluation of League roster rules and initiatives, assessing their impact on competitive performance, player investment, salaries, roster construction, player movement, and other relevant outcomes," and "leverage event, tracking, roster, financial, and other soccer-related data to evaluate player, club, and competition performance using advanced analytics, statistical methods, and data science techniques."
A closer look: how well can we trust the numbers?
Today, I want to talk about Daniel Job. Is he a good player?
That isn't an easy question to answer. One analyst could point to Job's 7 goals and 8 assists this season prior to transferring to FC Dallas. The analyst could tout his 93rd percentile offensive Impect packing scores and say "this kid is 21, he's only going to get better." Another analyst could argue that the fact these numbers came in the Norwegian second division render them useless. They could point to Job's meager two goals in the first division and claim that there's little chance he'll perform well in a superior league. Both takes miss context, but quantifying the level of context we're missing is a difficult task.
The reason I chose Job to illustrate this example is two-fold. First, this bluesky post from Tim Keech:
I'll get to the second reason in a second. This post struck me because it reminded me of the André Luiz dilemma — that is, how do we a judge a player with limited minutes who shows promise but doesn't have a large enough sample size? Both Job's $5m deal and Luiz's $18m move to SKC (as a DP!) are risky moves for each club. Luiz's 18th minute stunner for KC in their 3-1 win over LAFC on Saturday is a promising sign that the investment wasn't misguided, but that's still just one goal. Job has yet to debut for Dallas since joining the club.
In thinking about both deals, I recalled a post from Ben Griffis on LinkedIn I'd seen a few months prior:
Turns out, Ben had already answered my general question and my specific one in one fell swoop. The general question was "how do we quantify our confidence in a player's numbers?" while the specific one was "is Daniel Job good?"
Ben offers a Bayesian approach to measuring percentile scores and uses that to come up with two new measures: validity — quantifying trust — and reliability — signal vs noise. One of the most interesting parts of Ben's analysis is that his reliability ceiling (based on historic values in the same league and position) is 54%. From Ben:
These ceilings are low because football data, especially on a match-by-match level, can be extremely noisy. Trying to quantify that is an important step in the decision-making process, either as a scout recommending a player get more attention, or in a final transfer decision.
Daniel Job's 93rd percentile offensive packing had a reliability of just 30% after 11 matches. I'd be interested to see what it looked like over 18 matches (my guess is it would be higher, but given his past performances, I don't imagine it reaching that ceiling). In the end, that whichever player a club will sign has a 50/50 chance of replicating their previous numbers is both frustrating and logical. My curiosity leads me towards how these kinds of models can be built to incorporate league adjustments, too. How do Job's numbers in the second division versus the first division impact that score? Can we use league adjustments as weight factors in the model rather than multipliers, as they're often used, and can we build similar models to quantify certainty in those league translation estimates?
I foresee a workflow that looks something like this: Daniel Job, a player in the Norwegian second division, has a prior equivalent to that of an average 20-year-old winger in the Norwegian first division. He has some promising minutes in the first division, but the low quantity hardly moves the needle from that prior. Then, in the second division, he breaks out, but because our league translation measure (which has high confidence between the first and second divisions) suggests the second division is, say, 80% of the first division, we include that weight when calculating the adjustment. His value measure has moved up, but less than if he performed at that same level in the first division. FC Dallas look to sign Job, and consider how he would translate to MLS. Dallas staffers estimate that MLS is 200% of the Norwegian first division, but there isn't a lot of confidence in that measure since there haven't been a whole ton of players who have played in both sequentially. So, when predicting Job's future performances, the club shows a wide band of possible outcomes for that translation. They could also estimate his future performances within the Norwegian first/second division, and based on this hypothetical, formulate a potential value predicted from previous Norwegian breakouts with similar stats at the same position. Using that hypothetical market value, and the marginal returns of Job's potential production at the MLS level, FC Dallas can now compare him to options within the league (using a similar process, where the future performances of MLS-proven players are more certain) to determine whether the risk is worth it when compared to what talent the club could acquire within the league for an equivalent price.
Ultimately, the above optimistic workflow is simply a more quantified approach to thought processes that are surely occurring in every sporting director's brain. But what is analytics if not the attempt to quantify what we can already see with our own eyes and analyze within our brains?
If anyone has any further reading on bayesian uses in soccer — primarily how it can be used for player evaluation and/or league translation — please send it my way.