researchers.one - Harry Crane
The Kelly criterion prescribes the stake that maximizes the long run growth rate of bankroll when the bettor’s edge is known. In most practical situations, the edge is not known, but estimated. Based on the estimated edge, many practitioners opt to stake a fraction of the amount prescribed by the Kelly criterion. Common choices for fractional Kelly range from 0.10 to 0.50 of the full Kelly stake, with the choice guided primarily by folklore and heuristics. We reframe bet sizing as a constrained optimization depending on two parameters the bettor controls: (i) a stop loss E, the fraction of bankroll the bettor is willing to lose before abandoning the strategy, and (ii) a premature stopping probability π, the probability of hitting the stop loss even though the assumed edge is correct. We show that when the edge is accurately estimated, the probability of hitting the stop loss is (1−E)^θ∗ , where θ∗ is the Cram´er root of the per-round log-return. The constraint that this probability not exceed π determines the optimal stake in closed form: expressed as a fraction of the Kelly stake, α = 2 log(1−E) / (log(1−E) + log π) in the small-edge limit. Closed forms follow for the probability of reaching a target multiple of bankroll before the stop loss and for the degradation of the guarantee under an overestimated edge together with its robust correction.
substack.com - Alex Marin Felices
Separating individual football skill from team strength using a Bayesian Soccer Factor Model.
github.com
We are your essential toolkit for computer vision. From data loading to real-time zone counting, we provide the building blocks so you can focus on building applications around your models. 🤝
pysport.org
MORE DATA. MORE COMMUNITY. MORE IMPACT.
The first Analytics Cup brought the open-source football analytics community together to explore SkillCorner data, develop original ideas and present six outstanding projects at the final in Paris.
The winning project, Amar Shah's Simulated Annealing for Positional Optimisation, has since been developed into a module within the DataballPy Python package, making the work available for the wider analytics community to use and build upon.
Now the challenge returns as Analytics Cup 2.0. This second edition introduces a new Basketball dataset alongside Football, plus separate Europe and USA competitions designed to bring more of the global sports analytics community together. Both competitions will culminate in a live regional final in February 2027.
readthedocs.io
Welcome to the official documentation of DataBallPy!
This package is developed to create a standardized way to analyse soccer games using both event- and tracking data. Other packages, like kloppy and floodlight, already standardize the import of data sources. The current package goes a step further in combining different data streams from the same game. In this case, the Game object combines information from the event and tracking data. The main current feature is the smart synchronization of the tracking and event data. We utilize the Needleman-Wunch algorithm, inspired by this article, to align the tracking and even data, while ensuring the order of the events, something that is not done when only using (different) cost functions.
Although reading in and synchronising data is already very helpfull to get started with your analysis, it’s only the first step. Even after this first step, getting your first ‘simple’ metrics out of the data might be more difficult than anticipated. Therefore, the primary end goal for this package is to create a space where (scientific) soccer metrics are implemented and can be used in a few lines. We even plan to go further and show clear notebooks (to combine text and code) with visualizations for all the features we implement. This way, you will not only get easy access to the features/metrics, but also understand exactly how it is calculated. We hope this will inspire others (both developers and scientist) to further improve the current features, and come up with valuable new ones.
arxiv.org - Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David Holzmüller, Alan Arazi, Alexander Pfefferle..
Abstract:Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry. Research communities across disciplines are increasingly evaluating tabular foundation models on diverse datasets and tasks. However, these task- and discipline-specific evaluations remain largely inaccessible to model researchers because benchmark software and evaluation protocols are fragmented. As a result, model researchers rely on standard benchmarks, which are mostly defined for tasks where tabular foundation models already excel. The most challenging scenarios are excluded, limiting meaningful progress in the field by focusing on marginal improvements on IID data rather than on broader, more demanding challenges. To overcome this, we introduce BeyondArena, the first unified holistic benchmark for tabular data that supports diverse task types (IID, temporal, grouped), across sample size and feature dimensionality scales, with diverse feature types (with text, with high cardinality) from a broad range of disciplines. To enable unified benchmarking beyond standard benchmarks, we introduce Data Foundry, a Python framework and metadata schema for curating tabular datasets for predictive machine learning. Our results across 11 models and 142 curated datasets show that existing tabular foundation models excel on tiny- to medium-sized IID data, while traditional tree-based and deep learning models still dominate on non-IID, large, and high-dimensional datasets. BeyondArena guides model research for the most demanding challenges in tabular data, enabling progress towards truly foundational tabular models.
ssrn.com - Giovanni Angelini, Luca De Angelis, Carl Singleton
Studies of financial market informational efficiency have proven burdensome in practice, because it is difficult to pinpoint when news breaks and is known by some or all the participants. We overcome this by designing a framework to detect mispricing, test informational efficiency and evaluate the behavioural biases within high-frequency prediction markets. We demonstrate this using betting exchange data for association football, exploiting the moment when the first goal is scored in a match as major news that breaks cleanly. There are pre-match and in-play mispricing and inefficiency in these markets, explained by reverse favourite-longshot bias (favourite bias). The mispricing tends to increase when the major news is a surprise, such as a goal scored by a longshot team late in a match, with the market underestimating their chances of going on to win. These results suggest that, even in prediction markets with large crowds of participants trading state-contingent claims, significant informational inefficiency and behavioural biases can be reflected in prices.
