Initialising...

Explore the Model

Click on the pitch to plot a shot and get its predicted xG value.

From Raw Data to Probabilistic xG Mapping

What is Expected Goals (xG)?

Expected Goals, or xG, is a metric in football analytics that quantifies the quality of a goal-scoring opportunity. Rather than simply counting shots, xG assigns a probability to each attempt, representing how likely it was to result in a goal. A shot with an xG of 0.1 is expected to be scored 10% of the time, whilst a close-range tap-in might carry an xG of 0.9.

This enables a more nuanced analysis of team and player performance, looking beyond the luck often inherent in a final scoreline. A team that consistently creates high-quality chances is likely performing well even when the goals are not coming. This interactive plotter is built on a bespoke xG model trained from the ground up.

Step 1: Gathering the Data

Every good model starts with good data. The foundation of this project is a dataset of hundreds of thousands of shots, scraped from the public football analytics website Understat. A custom Python script gathers data across multiple seasons from Europe's top leagues, including the Premier League, La Liga and the Bundesliga. It extracts detailed shot-level data from match pages, capturing everything from pitch coordinates to game situation, and saves it for the next stage.

Step 2: Preparing the Data

Raw scraped data is rarely clean. A cleansing step handles inconsistencies before the dataset moves into feature engineering. The raw X and Y coordinates of a shot are useful on their own, but more powerful predictive features can be derived from them. Key engineered features include:

  • Distance to Goal: The Euclidean distance from the shot location to the centre of the goal.
  • Angle to Goal: The angle in radians that the shooter has on goal, calculated using vectors to each goalpost. A wider angle generally represents a better chance.

Categorical variables such as Situation (e.g. Open Play, Set Piece) and Shot Type (e.g. Head, Right Foot) are converted to a numerical format using one-hot encoding so the model can interpret them correctly.

Step 3: Training the Models

With a preprocessed dataset in place, the next step is training the predictive models. This project uses Logistic Regression, a robust and interpretable algorithm well-suited to binary classification tasks like predicting a goal (1) or no goal (0).

Rather than a single general-purpose model, four distinct models are trained to provide more specialised predictions. The correct model is selected automatically based on the inputs provided in the Controls panel:

  • Basic Model: Uses only location-based features (coordinates, distance and angle).
  • Situation Model: Adds the game situation (e.g. Open Play, Penalty).
  • Shot Type Model: Adds information about how the shot was taken (e.g. Head, Left Foot).
  • Advanced Model: The most comprehensive model, combining all available features for the most detailed predictions.

Each model is trained using scikit-learn, with hyperparameters tuned through randomised search and cross-validation.

Step 4: Pre-calculating Heatmaps

The Heatmaps page visualises xG values across the entire pitch. Computing these on demand for every visitor would be slow and wasteful, so a dedicated script (generate_heatmaps.py) runs as part of the data pipeline to pre-calculate them.

The script iterates over a fine grid of pitch coordinates and calculates the xG at each point for every combination of situation and shot type. The results are saved to a single JSON file. When a filter is selected on the Heatmaps page, the browser reads the relevant pre-calculated grid directly from that file with no server round-trip required.

Step 5: Exporting Models to ONNX for the Browser

The trained scikit-learn pipelines cannot run natively in a browser, so the final step of the data pipeline converts each one to the ONNX (Open Neural Network Exchange) format using skl2onnx. The resulting .onnx files are written directly into the frontend directory alongside the heatmap JSON.

The site originally used a Python Flask API hosted on Render to serve predictions. Render's free tier spins down inactive instances after a period of inactivity, which meant every first visit triggered a cold start that could take upwards of 60 seconds before the page became usable. Migrating to in-browser inference via ONNX Runtime Web eliminated that problem entirely. Predictions now run locally in the browser using the exported model files, with no server involved at all.

Step 6: Deployment

Because all inference happens client-side, the entire application is served as a static site via GitHub Pages. There is no backend to maintain or host. The ONNX model files and heatmap data are committed to the repository as static assets and served directly to the browser alongside the HTML, CSS and JavaScript.

Step 7: Explore the Code

The full project, from the data pipeline to the frontend, is open source. The repository includes the scraping scripts, feature engineering, model training, ONNX export pipeline and all frontend code.

View on GitHub