OnisAI · Shipped

Screenshot in, scored decision out

A multi-tenant operations platform for rideshare drivers. It ingests an offer as an image, parses it out of noisy OCR, scores it against a model of what the job actually pays, and carries it through its lifecycle with a full audit trail.

Software Developer & AI Automation · Contract Oct 2025 – Jun 2026 onisai.com Live demo GitHub See stack

A rideshare offer arrives on a countdown. The screen shows a fare and two distances, but never the number that decides whether the job is worth taking: what it pays per hour once the unpaid miles to reach the passenger and the fixed overhead of every job are taken out of it. There is not enough time to work that out by hand, so drivers fall back on rules of thumb that are wrong in both directions.

There is a working demo of the scoring engine at onisai.com/demo, and a longer technical writeup at onisai.com/engineering. The demo runs the production scoring module in the browser against synthetic offers.
One real offer, one tap, one explainable recommendation. A £19.78 offer arrives, the floating one-tap control runs the Shortcut, and the genuine result returns £0.55/minute and £1.92/mile. The scoring rules recommend ACCEPT: pay is GOOD and the pickup is CLOSE. The live wait is compressed for clarity. Exact route, addresses, postcodes, passenger rating and timestamps are intentionally blurred. OnisAI is an independent driver tool and is not affiliated with Uber.
Read the 26-second visual transcript
  1. 0:00–0:03: The driver waits in the normal app; the live wait is shown at 4×.
  2. 0:03–0:05: A £19.78 offer arrives with a 5-minute, 1.2-mile pickup and a 29-minute, 9.1-mile trip.
  3. 0:05–0:09: A coach mark identifies AssistiveTouch; one tap triggers the real capture flash and Shortcut progress.
  4. 0:09–0:11: The genuine scored notification returns while the offer remains visible.
  5. 0:11–0:15: £0.55/minute counts pickup, trip and two minutes of fixed overhead; the trip-only comparison is £0.68/minute.
  6. 0:15–0:19: £1.92/mile counts both pickup and trip distance; the trip-only comparison is £2.17/mile.
  7. 0:19–0:26: OnisAI recommends ACCEPT: the pay rule is above its £28/hour target and the pickup is close for a 29-minute trip. The driver still makes the final choice.

What I worked on

Built with a cross-functional team. My contribution was the OCR pipeline, the scoring engine, backend components for parsing, validation, summarisation and reporting, and dashboard UI work.

OCR ACCURACY · ~98%

Before

A misread postcode digit, a floating UI label bleeding into the address line, or an address split across two OCR lines — any one silently corrupts the trip record.

After

~98% of captured offers parse clean on the first pass, no human involved.

How: parseOcrText() runs postcode OCR-confusion correction, multi-candidate price selection, overlay-stopword trimming and cross-line address stitching before any field is trusted.

MANUAL PROCESSING · ~90% less

Before

An offer that failed to parse landed in a parse-error state — someone had to open the trip and type trip minutes, trip miles, pickup minutes, pickup miles and price by hand before it could be scored.

After

~90% fewer offers ever reach that manual-correction step.

How: the same parsing guardrails above catch what would otherwise fall through to a person.

END TO END · ~1.2s

Before

The driver does the pay-per-hour maths themselves, in their head, against a live countdown.

After

The score lands in about a second — screenshot in, verdict out, before the acceptance window closes.

How: computeOfferScores() is a pure function — no database or network calls in the maths itself; OCR and parsing are the only real work.

METRICS PER OFFER · 9

Before

The screen shows a fare and two distances — never the number that decides whether the job is worth taking.

After

One function call returns per-minute rate, per-mile rate, both again including pickup, nominal hourly, overhead-adjusted hourly, pay status, pickup status and fuel cost.

How: computeOfferScores() in src/scoring.js.

The OCR figure is the one that mattered commercially. A parser that is right most of the time is worse than useless when each mistake costs money, which is why the guardrails above and the manual-correction path both exist.

Three decisions worth explaining

Pickup cost is relative, not absolute

A two-mile pickup is wasteful for a three-minute hop and trivial for a forty-minute motorway run. So the thresholds are not fixed. The engine classifies the trip into a band first — short, medium or long — then judges the pickup against that band.

There is a deliberate asymmetry inside each band: a CLOSE verdict requires both distance and time to be within bounds, while SLIGHTLY FAR accepts either. Distance and time diverge under traffic, and a pickup that is short in miles but long in minutes is a traffic problem rather than a distance one. It should degrade to a warning, not pass as clean.

This is also the idea the corridor engine later generalises: once you accept that context changes what a number means, the next question is what the route between those two points passes through.

Charge the overhead before judging the rate

Nominal hourly rate pretends a job ends the instant the passenger steps out. Every job really costs a couple of minutes of waiting, loading and repositioning, and that overhead is proportionally brutal on short trips. Charging it first is what stops the engine over-rewarding quick hops that feel lucrative and are not.

scoring.js
hourly_adj = price / max(trip_min + overheadMinutes, 1) * 60

The overhead and both accept/reject thresholds are environment-configurable rather than hard-coded, because the right numbers differ by city, vehicle and driver. The defaults are starting points, not truths.

Degrade, do not crash

The capture path is the one thing that must never fail: an offer is gone in seconds and will not be shown again. So Cloud Vision is an optional dependency — required lazily, with the load error cached and a structured failure returned rather than thrown.

A missing OCR provider therefore degrades one request instead of taking down the process. A bad parse is recoverable too: the trip lands in offered_parse_failed with its raw OCR text intact, and the driver corrects the fields by hand. Those corrections are then rescored through the same function as the original parse rather than patched into the stored row — otherwise a trip can end up displaying a fare and a verdict that contradict each other.

An event log, not a status column

Trips move through offered, accepted and completed, with rejected and cancelled as terminal alternatives, and invalid transitions are rejected explicitly rather than silently succeeding. Every transition appends to an event log, and each event carries a snapshot of the prior state.

That one choice makes undo nearly free: the undo path reads the last status event, restores the previous state wholesale, and appends an event of its own. No reverse state machine to write and keep correct, and no history destroyed — the undo is itself a record.

It matters because a driver tapping the wrong button in a moving vehicle is not an edge case, it is routine. Undo had to be one tap and impossible to get wrong.

Closing the loop

The platform's quoted trip time is an estimate; what actually happened is knowable only afterwards, and the system holds both numbers. On completion it compares elapsed time against the original estimate and grades the difference into a traffic level.

The grading is not a plain threshold ladder. Long trips absorb small absolute overruns — three minutes late on a forty-minute journey is noise, not congestion — so any trip of twenty minutes or more landing within ±10% of estimate is treated as normal flow regardless of the absolute delta. Judging delay in raw minutes would systematically over-report congestion on exactly the long trips the engine most wants to recommend.

This is the feedback loop that made the corridor work possible: once every completed trip produces a data point about how a route behaved at a given hour, evidence-based routing stops being hypothetical.

Watching the engine work

One offer a second is easy to reason about by hand. Fifty is the point where a driver needs the engine, not a calculator. Below, fifty generated offers stream through the same rules documented above, one at a time, so the accept/reject calls and the pickup and traffic classifications are visible while they happen rather than as a single static result.

LIVE SCORING FEED · SYNTHETIC OFFERS, REAL RULES

Fifty generated offers run through the same pickup-distance, traffic and congestion-zone logic described above. The prices, distances and timestamps are synthetic; the classification rules are the production ones. The GOOD/BAD hourly split follows the £28/hour target from the demo video; the exact reject threshold shown here is an illustrative default, not the production config value.

0offers ingested
0%flagged GOOD
0bad trips caught
£0.00/hrrunning average, adjusted
#PriceTripPickupHourly (adj)VerdictPickup statusTraffic

What the same offers teach the pricing memory

The scoring table above is the whole story in production — it decides accept or reject on its own. But every completed trip also feeds a separate statistical layer, src/self-learning-engine.js, that never overrides the rules above; it just quietly builds evidence for whether it could be trusted to one day. It keeps a running mean and variance (via Welford's online algorithm) of clean rate per minute for every geography × day-type × time-of-day context, plus a separate memory of specific pickup→dropoff route edges, and scores its own readiness to graduate from “shadow mode” toward guiding real decisions. The panel below runs that exact math against the same fifty offers, live.

PRICING MEMORY · WELFORD'S RUNNING AVERAGE

Each card is a running £/hour mean for one context, computed exactly as updatePricingBucket() does in production. Confidence follows the real formula too: it grows with sample count and shrinks with variance, never just with time elapsed.

ROUTE MEMORY · PICKUP→DROPOFF EDGES

    Learning-replacement readiness 0%

    Keep the deterministic system primary while the learner builds memory.

    This is the pricing half of buildLearningReplacementReadinessFromTenantState(), scored live against these fifty offers. The parser-repair half isn't shown here since this browser demo has no OCR correction memory for it to learn from.

    Node.js Express PostgreSQL Redis Nginx PM2 Stripe Tesseract.js Google Cloud Vision Claude API Codex Web Push / VAPID

    The real production stack: a Node.js/Express backend on PostgreSQL and Redis, deployed with PM2 behind Nginx on a Hetzner VPS. Billing runs on Stripe (Checkout, Billing Portal, webhook sync); OCR falls back between Tesseract.js and Google Cloud Vision depending on config; and an optional layer calls Claude directly (no SDK, raw /v1/messages calls) for admin-facing trip summaries and as a secondary extraction/reconciliation check alongside the deterministic parser above.

    The build process is disclosed too: this repo is developed under a Codex + Claude dual-agent workflow with a shared quality gate, documented in its own AGENTS.md and CLAUDE.md — not solo-coded.