Skip to content
AI & Model Training

Train models on the official esports dataset.

Six titles. Multi-year history. Billions of structured events. All rights-cleared for commercial AI training and live inference - the foundation your model legal team can sign off on.

Sample event · CS2
PLAYER_DIED
{
  "event": "PLAYER_DIED",
  "game_time": 142.318,
  "round": 14,
  "map": "de_inferno",
  "actor": { "id": "s1mple", "team": "NAVI" },
  "weapon": "awp",
  "headshot": true,
  "economy": { "ct_eq": 22400, "t_eq": 14250 },
  "win_prob_at_event": 0.71,
  "source": "official_game_server",
  "license": "commercial_ai_ok"
}
~1.27M structured events like this per match.
Illustrative schema · Production schema documented in the API spec.
230k+
Official matches in the historical archive
5+
Years of structured history across titles
100%
Rights-cleared for commercial AI use
The problem

Esports AI is bottlenecked by the dataset, not the model.

1

Public datasets

Narrow, outdated, often single-title hobby drops. Fine for a Kaggle notebook. Untenable for a production model.

2

Scraped data

Lossy, unofficial, and legally exposed for commercial use. Schema breaks every time a game patches.

3

Build-your-own pipelines

A multi-year engineering project. New title, new pipeline. Coverage gaps you'll find in production.

So the question is - what does it actually take to train models that ship?

The platform

GRID is the official esports data platform.

Captured directly from the game server, under rights agreements with the publishers and tournament organisers who own the data. One schema. Every covered title. Licensed for commercial AI.

Official data partnerships - direct from rights holders
Riot Games
ESL FACEIT
BLAST
Ubisoft
Moonton
Esports World Cup
+ 70 TOs
The dataset

The largest official esports training corpus.

Six flagship titles. Years of history. Live, real-time events on top.

6
Flagship titles, one schema
230k+
Matches in the historical archive
1.27M
Structured events per match (avg.)
5+
Years of structured history
What you get

Official. Structured. Rich enough to ship.

01Official & licensed

Sign-off ready.

Rights-cleared for commercial AI training and live inference. The dataset your legal team will actually approve - and your model can ship on.

100% licensed for commercial AI.
02Structured & standardised

One schema. Every title.

A single, versioned data model across every covered title. No per-title cleaning, no per-update breakage, no schema drift mid-training run.

v3 stable, versioned schema.
03Rich & granular

Server-level depth.

Player events, economy state, positional data, objectives. The feature richness models need - not just box scores.

1.27M data points per match (avg.).
Delivery formats

Train, infer, backtest. Same data, your pipeline.

REST API

Pull on demand

Hit any match, any event window. Good for offline training, evaluation, and on-demand inference.

WebSocket

Stream live

Sub-200ms live match state for real-time inference, agents, copilots, and live decisioning.

Bulk export

Train at scale

Multi-year archives delivered as Parquet, JSON, or CSV. Built for training pipelines and warehouses.

Snapshots

Backtest cleanly

Point-in-time match state for backtesting and reproducibility. Run the same eval on the same state every time.

What you can train

Models that ship, on data they can trust.

Win probability models

Forecast match, map, and round outcomes from live game state - for markets, content, or engagement.

Player & team performance

Predict KDA, contribution, form, and matchup effects across players, teams, and series.

Markets & odds models

Sharpen pricing, settlement, and risk models on the same official data settlement is run against.

AI agents & copilots

Build esports-aware assistants, fan chatbots, content generators, and broadcast copilots on live data.

Computer vision

Train and validate vision models against ground-truth official data - events, positions, outcomes.

Integrity & anomaly detection

Identify match-fixing patterns, behavioural outliers, and integrity flags at population scale.

Integration

One source. Every delivery path.

  1. Source
    Game server
  2. Platform
    GRID Data Platform
  3. Delivery
    REST · WebSocket · Bulk · Snapshots
  4. Destination
    Your pipeline · Train · Eval · Infer
Single versioned schema · Same source for training and inference
FAQ

The short answers.

Is GRID data licensed for commercial AI training?

Yes. The commercial tier grants rights for training and live inference, including for production AI products. For non-commercial research or student projects, get in touch about the right terms.

Which esports titles are in the dataset?

Six flagship titles - League of Legends, VALORANT, Counter-Strike 2, Dota 2, Tom Clancy’s Rainbow Six Siege, and Mobile Legends: Bang Bang - normalised into one consistent schema. New titles ship through the same data model.

How much historical data is available?

230k+ official matches in the historical archive across covered titles, with 5+ years of structured history. Roughly 1.27M structured events per match on average.

What delivery formats do you support?

REST for on-demand pulls, WebSocket for live streaming, bulk exports (Parquet, JSON, CSV) for training pipelines, and point-in-time snapshots for reproducible backtesting.

How do you handle schema changes when games update?

GRID maintains a stable, versioned schema across game updates. Breaking changes are flagged ahead of time with a version bump and migration notes - so a mid-training game patch doesn't break your pipeline.

How do I get started?

Tell us about your project and the GRID team will scope the right access, from research and prototyping through to production. Get in touch to start.

Get in touch

Ship esports AI on official data.

Whether you're training a production model, building an esports-aware agent, or running live inference - let's talk about what you're building.

  • Dataset walkthrough & schema sample
  • Commercial licensing & SLAs
  • Response within 24 hours