Walk into any gym, point your camera at an unfamiliar machine, and keep moving. GymLens removes the search, translates the equipment into clear exercise options, and carries your choice straight into a live workout.
ENGINEERING CASE STUDY / 78 SECONDS / REAL USER JOURNEY
From an unknown machine to a bounded decision
The problem, the product response, and the system behaviour behind one photograph.
Read the 78-second visual transcript
- 0:00–0:09: A real gym moment establishes the bottleneck: the machine is visible, but the user does not know the words needed to search for it.
- 0:09–0:17: The category workflow is mapped as three dependencies: name the machine, identify the movement, then find the equipment category.
- 0:17–0:27: The GymLens journey begins: a person photographs the machine while the real product interface shows how physical evidence enters the app.
- 0:27–0:37: The capture boundary validates image type and size, refusing an invalid request before it can consume inference resources.
- 0:37–0:50: A bounded image request moves through inference across 25 generic classes, becomes ranked candidates, and reaches human confirmation.
- 0:50–1:01: The trust and privacy boundary separates the confirmed label from the photograph, which is processed and discarded.
- 1:01–1:18: The confirmed exercise enters the workout, becomes a local record, and produces visible progress.
Built beyond the demo
GymLens is a complete training workflow, not a camera experiment: a broad machine catalogue, exercise discovery, workout tracking, history, performance summaries and a tested delivery pipeline.
Two runtimes, one product boundary
The public PWA is intentionally inexpensive to operate: static assets at the edge, a Worker for recognition requests, Workers AI for hosted inference and D1 for confirmations. Workout profile, sessions and history remain in browser storage.
Camera, catalogue, workout timer, history and performance UI.
Validates uploads, calls hosted inference and returns bounded suggestions.
Processes pixels, discards the photo and stores the user-confirmed label.
PostgreSQL, SigLIP experiments, fixtures and regression evaluation in Docker.
The engineering behind the experience
The model suggests; the person decides
Gym equipment changes colour, frame geometry and branding without changing its purpose. The recognition target is therefore a generic machine class, not a specific product. Results are always confirmable or correctable because confidence is not the same as accuracy.
recognise(image_bytes, known_machines) -> list[Candidate]
# The endpoint depends on the contract, not a specific model.
# CI uses a deterministic adapter; research can swap in SigLIP.Keep policy in one place
An early build allowed a confidence threshold to drift between two files. Both implementations looked reasonable and all local tests passed. The durable fix was one guardrails module that keeps prose, thresholds and boundary tests together.
if result.requires_confirmation:
show_ranked_suggestions()
else:
refuse_to_guess()Store less by default
Gym photographs may include other people. The hosted app processes the upload for recognition and does not retain it. It keeps only the information needed for the product feedback loop, while workout history remains local to the browser.
Designed to improve with every confirmation
The recognition layer and product interface are deliberately separated, so the model can improve without rebuilding the workout experience. Every confirmed or corrected result creates better evidence for evaluation while the user stays in control of the session.