Data quality
A blank in this database is an unestablished fact, never a zero. These are the gaps that are known and stated.
Known gaps
fastest_lap — the fastest lap of a race the pole harvest has not reached yet
The fastest lap of every race comes from harvest/poles.txt, which is written by hand. Everything else about a completed race - the classification, the qualifying sheet, the standings, and since this version the driver at the front of the grid - is refreshed from F1DB by a scheduled job within a day or two of the flag. So for the week between a Grand Prix and somebody editing that file, the site shows a completed race with no fastest lap. Pole used to sit in the same hole and no longer does: F1DB may now fill grid 1 where nothing else has.
F1DB publishes the fastest lap per race in its own file, which tools/f1db_fetch.py does not read - it takes race-results.yml and not the fastest-lap results beside it. Adding that reader, a harvest/fastest_laps.txt beside the others, and a loader that fills race_entries.fastest_lap only where the harvest has not, closes this the same way pole was closed. Until then verify.py warns for the current season, which is the signal that the harvest is behind.
finish_position — shared drives, and where two sources read a race differently
CLOSED in v2.15, and by a licence rather than a harvest. The full classification - 27,555 entries across all 1,161 races, 1950 to 2026 - now ships in the committed database. It comes from F1DB, which is CC BY 4.0: attribution only. The same facts via Jolpica-F1 carry Ergast's CC BY-NC-SA, and that non-commercial clause is the whole reason this gap stood for seven versions; nothing about the data was ever hard to get. The winner of every one of the 1,161 races was already held here from the Wikipedia harvest, and F1DB agreed with all of them. What remains is not missing data. race_entries is ONE ROW PER DRIVER PER RACE, so a driver who drove two cars in one Grand Prix - normal before 1965, and the 1955 Argentine Grand Prix in particular - can only keep one result. And the two sources genuinely differ on 118 of 26,082 entries when Jolpica is loaded on top: F1DB leaves a disqualified driver's position VACANT while Jolpica promotes everyone below, so the 1983 Brazilian Grand Prix has no second place in one reading and Lauda second in the other. Neither is wrong.
Nothing to fetch. Run tools/ergast_load.py --from-dump to put Jolpica alongside: it no longer overwrites, it records every disagreement in `discrepancies` and leaves the stored value alone. Reading those 118 rows is the work, and it is a person's.
chassis_id — the chassis each race was won in, where a season is ambiguous
The winning chassis is known for 874 of 1,161 races. It comes from F1DB's entry lists, resolved through the driver and the round: F1DB records which chassis a constructor ran in a SEASON and which driver an entrant ran in which ROUND, and the second is far sharper than the first. What is left is the entrants that ran more than one design and do not say which raced where - Gold Leaf Team Lotus in 1970 entered a 49C, a 72B and a 72C, so Rindt's wins stay NULL while Soler-Roig's single Garvey Team Lotus entry resolves. v_ambiguous_seasons lists them. The gap is concentrated before 1980 because a modern team runs one car all season while a 1960s constructor was a name several privateers entered several different chassis under.
Per-round chassis data, which no source in use here has. The season articles print a chassis column in the same results table the winners came from, and that column IS per round. Harvesting it is only safe with the check that now exists: a harvested chassis must appear in that constructor's entry list for that season, or be refused.
cars.poles — the car each pole was taken in, where the season is ambiguous
This was the largest gap in the car data and is now mostly closed. The pole and fastest-lap harvest recorded who set them but not what they drove, so 1,260 of 2,424 entries carried no constructor at all and could not reach a car. F1DB's per-round driver data supplies it: 99 per cent of entries now name a constructor, and NINE cars match their published career pole total exactly, where previously none could be checked for more than not exceeding it. What is left is the seasons a constructor ran more than one chassis, where attributing a pole to a particular car would be a guess - McLaren ran the M23 and the M26 through 1976 and 1977, and the blanket season claim gave the M23 sixteen poles against a published fourteen before this was checked.
Per-round chassis data, which no source in use here has. The remaining seasons are listed in v_ambiguous_seasons and in car_seasons where corroborated = 0.
laps — lap times, tyre stints, pit stops, radio and telemetry
The laps, stints, pit_stops, race_control_messages and team_radio tables are EMPTY in the distributed database, and that is a licensing and reproducibility decision rather than a missing harvest. Two sources fill them, and between them they reach further back than the 2018 floor this gap used to describe. Jolpica's database dump has 628,454 race laps from 1996 and 12,627 pit stops from 2011 - twenty-two seasons more than was previously thought available - but its Ergast lineage is CC BY-NC-SA, non-commercial, so the rows are loaded locally and never committed. FastF1 reads the F1 live timing API, which begins in 2018 and for which nothing earlier exists in any retrievable form; it adds sectors, tyre compound, stints and race control, none of which Jolpica has. Both may hold the same race at once - the schema keys a lap on (race, source, driver, lap) precisely so they can be compared rather than one overwriting the other. Car telemetry (speed, throttle, brake, gear at about 4 Hz) is deliberately NOT stored here at all: it is hundreds of megabytes per weekend and fastf1_load.py writes it to Parquet beside the database.
LOCALLY ONLY, and there is no version of this that ends in a shipped table: no source has lap times under a licence that permits passing them on. F1DB is the one source here that does permit it - which is why 22,481 of its pit stops ARE committed - and it has no lap times. So this is not an unfinished harvest, it is the correct state until that changes. On your own machine: python3 tools/ergast_load.py --from-dump --timing (1996-, one hash-verified zip, about fifteen seconds), and/or pip install fastf1 && python3 tools/fastf1_load.py --years 2018-2026 --results --radio. Then ./f1 laps. verify.py re-derives each race's fastest lap from the lap times and checks it against the setter already stored from the pole harvest - 446 races, no disagreement. docs/TIMING-ARCHITECTURE.md has the measurements and the design this would take if a redistributable source ever appears.
race_timing — pole, fastest lap and race times per race
The race_timing table is empty. These figures are published per race rather than per season, so filling them for 1950-2017 means reading 1,161 individual race articles, which was judged too expensive against what it returns. From 2018 the FastF1 loader supplies the same information at far greater resolution.
Either harvest race articles for the pre-2018 seasons, or accept that timing starts in 2018 and fill it with tools/fastf1_load.py.
layout_name — circuit configuration as raced, for most circuits
Thirteen circuits have a complete configuration timeline: Spa, Monza, Silverstone, Hockenheim, Interlagos, Indianapolis, Catalunya, Albert Park, the Spielberg site, the Mexico City site, Bahrain, Marina Bay and Yas Marina. For every other circuit the database holds one set of figures - the most recent - so a 1976 Kyalami lap is reported at the length of the 1992 rebuild. v_race_venues is explicit about this: its 'figures' column reads 'as raced' where a layout row covers the year and 'current layout' where it does not. Zandvoort, Suzuka, Imola, Jerez, Estoril, Paul Ricard, Zolder, Brands Hatch, Kyalami and Buenos Aires all changed shape between championship races and are not yet broken out.
Research the configuration history of each remaining multi-race circuit and add the rows; verify.py already enforces that a circuit's layout rows, once present, form a complete non-overlapping timeline.
fastest_lap — fastest lap, 2021 round 12 (Belgian Grand Prix)
No fastest lap is recorded because none was set. The race was abandoned after two laps behind the safety car, half points were awarded and no racing lap was completed. This is a true null, not missing data.
Nothing to fix - the absence is correct.
qualifying.q1 — qualifying session detail before 1996, and sector times
Qualifying is held for all 1,161 races - 26,975 rows - but its SHAPE changes. Before the knockout format a session is a single time, so q1, q2 and q3 are NULL and there is nothing to put in them; from 1996 the three segments are recorded and `time` is NULL instead. Neither is back-filled from the other, and verify.py fails the build if a row ever carries both. What is missing throughout is sector times, tyre compound and the lap a time was set on, none of which was published before the live timing era.
From 2018, tools/fastf1_load.py has all of it at far greater resolution. Before that it does not exist in any retrievable form.
centreline — the shape of a circuit, for anything but the present day
circuit_geometry traces a circuit from OpenStreetMap and checks the trace against the length this database already held. It can only ever be the CURRENT configuration, because OSM maps what is on the ground: Spa's 14.1 km Ardennes road course, Monza's banked sopraelevata and the 1976 Kyalami are not mapped and cannot be. Wikidata does model historic layouts as their own entities - 'Circuit de Monaco Grand Prix Circuit (1929-1972)' is Q66712049 - but those entities carry a length and a date range and NO coordinates, so there is no geometry source for them anywhere. A trace is therefore attached to a layout only where that layout is still current, and historic layouts have no row rather than a modern shape standing in for them.
Nothing available. A historic centreline would have to be traced from period maps or aerial survey, which is a research project rather than a harvest, and any such trace would have no independent length to be checked against - the one thing that makes the current ones trustworthy. Leaving them absent is the correct answer.
article_images.name_matches — whether a photograph shows the car
602 car articles carry a lead photograph from Wikimedia Commons, with its licence and photographer. The ARTICLE is well constrained - it passed the constructor, seasons and name checks in tools/wikispec_fetch.py before it was accepted - so the recorded claim is 'the article proved to describe this chassis leads with this file'. What is NOT established is that the photograph shows the car. Nothing in this database constrains the content of an image and there is no second source to disagree with, which makes this the only part of the database with no cross-check available at all. Testing whether the file name mentions the chassis finds 265 of 602, because most correct images are filed under the driver - File:Jos_Verstappen_2000_Monza_(cropped).jpg really is an Arrows A21 - so the test cannot be a rule without discarding half the good rows. It is stored as name_matches and enforced nowhere. The failure it half-detects is real: the ATS D5 article leads with a photograph of officials and police.
A person looking. v_images_to_check lists the 337 whose file name does not name the car, worst first by how many chassis depend on the article. Every row sits at 'unverified' until then, which is where this database puts what it cannot prove.