THE WALKTHROUGH
Why it exists. How you use it. What the engine does.
FILM 01 / 68 SECONDS
How drivers create traffic evidence
Follow a driver through three bottlenecks, record the beacons, then use the evidence on the return journey.
The whole lifecycle. The map builds, traffic fills it, one driver runs A to B and is held up at three bottlenecks - a beacon dropped at each, where and when it happened. Other drivers report the same way until the map has enough evidence, and the return leg is routed around the delays instead of into them. Corridor Engine grew out of OnisAI, a private-hire tool, but a blocked road costs any driver the same time. Road geometry follows Bloomsbury streets; driver activity, beacon counts and route outcomes are illustrative. Map © OpenStreetMap contributors.
FILM 02 / 60 SECONDS
Is this taxi offer worth taking?
A phone offer, a route check, two traffic delays and a decision to decline.
A £13 offer advertises 24 minutes. Two existing traffic beacons on its route add an illustrative four and five minutes: 33 minutes in total, with the fare unchanged. The driver checks the evidence and declines. This demonstrates a possible Corridor Engine use case; delay figures are illustrative, not measured predictions. Map © OpenStreetMap contributors.
FILM 03 / ENGINE INTERNALS
From traffic reports to a route score
Watch the concept, then explore real report counts and the Python scoring rule below.
A conceptual walkthrough of spatial matching, time relevance and storage. The video’s weights and stored-record fields are illustrative; the charts below show the collected data and the actual Python shadow-scoring rule.
THE DATA BEHIND THE DEMO
Real reports. A visible calculation.
The saved archive shows when reports were collected. The scorer shows how a report contributes to a route check. Explore both below.
When were traffic reports logged?
398 reports across 27 reporting days in June and July 2026.
91.5% were logged between 10:00 and 19:00. The sample is concentrated in daytime hours. Zero late-night reports means no evidence in that bucket, not clear roads.
All 398 records are labelled traffic; this archive has no clear-road comparison sample. Hours use the recorded timestamps and the Python scorer’s five buckets. Counts are reports, not unique drivers or measured delays. Download aggregate data.
PYTHON SHADOW SCORER / INTERACTIVE EXAMPLE
What adds up to the score?
Change the example inputs to see the actual scoring rule. These inputs are synthetic; the formula follows the local Python implementation.
- Exact × 3 6
- Nearby × 1 1
- Time match × 2 2
- Repeated-sector bonus 0
6 + 1 + 2 + 0 = 9 points · RED
GREEN: 0–1 · AMBER: 2–4 · RED: 5+
These are rule thresholds, not calibrated probabilities or extra journey minutes.
This implementation checks a straight line, not a routed road path. Reports beyond 250 metres add nothing. Native MapKit routing remains planned; the score is used in shadow mode.
How this relates to the video
The film illustrates matching, time relevance and storage. Its fractional weights and stored-record fields are explanatory examples. This calculator reflects route_beacon_shadow.py: 3 × exact + nearby + 2 × time-matched + a possible 2-point sector bonus. The separate offline map builder uses four time buckets; this scorer uses five. Neither this archive nor the film establishes prediction accuracy or minutes saved.
Separate public geographic snapshot, published 23 August 2026 (not the July archive chart above). Cells holding fewer than three reports are excluded, which is why 27 reports are withheld. These counts describe the evidence available — they are not a measure of prediction accuracy or of city-wide coverage.
04 / CHECK THE OFFER
Check traffic delays
before accepting a trip.
Corridor Engine checks the proposed route against driver-reported traffic. It identifies bottlenecks along the way and nearby evidence that may affect the journey, helping a driver assess whether an offer is worth the time.
The current engine records scores in shadow mode. The walkthrough illustrates an offer-checking interface; native iOS and MapKit integration are planned.
Explore the engineering details
The problem with endpoint scoring
The first version of this scoring worked on endpoints. Take the fare, the trip distance, the pickup distance, the destination postcode family, and decide. It is the obvious model and it is what most offer calculators do.
It is also wrong in a specific, repeatable way, and the failure is invisible in the data it looks at:
Endpoints describe where a job starts and finishes. They say nothing about the forty minutes in between. Two trips with identical fares, identical distances and identical destination postcodes can differ by twenty minutes of actual driving because one of them goes through a corridor that reliably jams at that hour.
So the engine had to answer a different question: what does this route actually pass through?
Evidence, not assumptions
The tempting fix is a congestion API or a static zone map. Both were rejected. A generic congestion feed describes traffic in the abstract; it does not know that one particular pull-out is bad specifically for a driver trying to turn right at 5pm. Static zone heuristics are worse — they are guesses with the confidence of rules.
Instead the system collects its own evidence. A second iOS Shortcut, deliberately kept to
two taps so it can be used while working, captures a GPS snapshot and records it as either
a traffic or no_traffic observation. Each
event is reverse-geocoded, reduced to postcode, outcode and sector, and stamped with a time
bucket.
| Bucket | Hours |
|---|---|
| night | 00:00 – 06:00 |
| morning | 06:00 – 11:00 |
| midday | 11:00 – 16:00 |
| evening | 16:00 – 24:00 |
Four buckets rather than twenty-four hours, because the database has to become useful early. Hourly resolution spreads sparse observations so thin that nothing reaches significance for months. Coarse buckets mean a handful of evenings in one area is already a usable signal, and the resolution can always be refined once volume justifies it.
Building the graph
Individual points are not corridors. To get from scattered observations to something with shape, consecutive events are joined into edges — but only when the join is plausible.
NODE_GRID_PX = 22.0
MAX_GAP_MINUTES = 35.0
MAX_EDGE_DISTANCE_METERS = 4800.0
MIN_EDGE_DISTANCE_METERS = 12.0
Each constant exists to reject a specific kind of false corridor:
| Constraint | What it rejects |
|---|---|
| gap ≤ 35 min | Two observations far apart in time. The driver went home in between; there is no corridor there. |
| distance ≥ 12 m | Repeated logs from the same spot while stationary, which would otherwise pile up as a fake dense node. |
| distance ≤ 4800 m | Jumps too long to represent continuous driving through observed conditions. |
| 22 px node grid | Snaps nearby points together, so one junction visited fifty times is one node rather than fifty. |
Without these, the graph connects everything to everything and confidently describes corridors that were never driven. The constraints are the difference between a map of observed conditions and an attractive-looking fiction.
Each node and edge carries per-bucket counts, and a dominant bucket is derived from them, so the graph encodes not just where conditions were observed but when.
Intersecting a route
With the graph built, scoring a live offer means constructing the corridor between pickup and dropoff and asking what falls inside it.
The preferred path takes a routed polyline from MapKit, buffers it into a corridor of configurable width, and intersects the beacon geometry against that shape. Where no routed path is available it falls back to a straight line between the endpoints, buffered wider to compensate for the fact that a straight line is a poor model of a road.
The corridor is never a zero-width line. Width absorbs route uncertainty, lane spread, and congestion bleeding outward from the road that causes it. The goal is operational usefulness rather than geometric purity — a mathematically exact line through a city is precisely wrong.
READING THE EVIDENCE
What each hit tells us
| Hit | Meaning |
|---|---|
| Exact | The beacon lies inside the primary corridor. |
| Near | Just outside, but close enough to suggest route pressure. |
| Timed | A hit whose time profile matches the current trip — same 15-minute window, weekday, or weekpart. |
Timed hits carry more weight than generic historical ones, which is the whole point of bucketing the evidence in the first place. A corridor that jams every weekday evening is irrelevant to a Sunday morning trip, and a model that cannot express that difference will keep declining good work.
The weighting runs in this order:
- exact hits that are also timed — strongest
- exact hits, untimed
- near hits that are timed
- near hits, untimed — weakest
Where the operator overrules the model
One deliberate design decision: explicit operator rules are evaluated separately from beacon evidence, and can override it.
Some corridors are known bad by direct experience long before the statistics agree. The Parkhurst Road and Holloway Road pull is a confirmed trap because the driver has repeatedly lived it. Waiting for the beacon database to reach significance before acting on that would be pedantry — discarding the most reliable evidence available on the grounds that it is not yet numerous.
Keeping operator hits in a separate field also keeps the decision explainable. The output
can say RED x2 and Parkhurst trap as
distinct reasons, rather than collapsing both into an opaque score. A driver who does not
understand why a job was flagged will override the system, and a system that gets
overridden is not deployed.
The output contract
The engine returns compact evidence rather than a bare verdict, so the decision layer can weigh it against fare, pickup cost and rating rather than being dictated to:
struct RouteCorridorEngineOutput {
let routeMode: RouteMode
let corridorWidthMeters: Double
let exactHits: Int
let nearHits: Int
let timedHits: Int
let weightedTrapScore: Double
let operatorRuleHits: Int
let matchedBeaconIDs: [UUID]
let matchedOutcodes: [String]
let matchedSectors: [String]
let primaryReason: String?
let confidence: Double
}
primaryReason and confidence are there for
the same reason as the separate operator field. A score with no explanation and no stated
uncertainty cannot be debugged, cannot be trusted, and cannot be improved.
What is proven, and what is not
Being precise about this matters more than making the project sound finished.
| Component | Status |
|---|---|
| GPS beacon logging from a two-tap Shortcut | Proven in live use |
| Postcode / outcode / sector extraction from noisy OCR | Proven in live use |
| Traffic evidence grouped by area and time bucket | Proven in live use |
| Corridor graph construction and visualisation | Built and running offline |
| Route-line scoring against beacon geometry | Shadow mode — scored but not acted on |
| Native iOS engine, MapKit routed corridors | Specified, not built |
Shadow mode means the corridor score is computed alongside real decisions and recorded, without influencing them. It is the only honest way to evaluate a scoring change: you find out whether it would have been right before you let it cost anyone money.
The native engine is a written contract rather than a shipped app — input and output shapes, route modes, weighting direction and failure behaviour. Writing that down before building it is how the Python implementation and the eventual Swift one stay the same product rather than two divergent guesses.
The live piece is a Pythonista/iOS Shortcuts workflow that logs GPS-tagged traffic beacons in the field and builds the grid-cell corridor graph from them. MapKit route-polyline matching and the SwiftData model above are written contracts for the native app this is designed to graduate into — not yet shipped code, which is why the chip above says “contracts”, not “MapKit”. The script also self-updates: it checks a JSON manifest served straight from a GitHub raw URL and pulls down new code automatically, the same delivery mechanism as the Wally Driving automation project. Like the other projects here, it was built with Claude and Codex rather than solo-coded.