FitGuru: An on‑device squat coach that counts reps, corrects form, and never uploads your video

Table of Contents

  1. Key Highlights
  2. Introduction
  3. Why “two brains” beats “one big model”
  4. From camera to meaningful joint angles
  5. A rep is a state machine, not a lone frame
  6. Gym Shield: locking onto the right person
  7. Local‑first data: privacy by design
  8. Post‑set cognition: Gemini 2.5 Flash and structured debriefs
  9. Giving the coach a voice: ElevenLabs, caching, and graceful fallback
  10. Hands‑free controls tuned for the gym
  11. Heart rate as a second sense
  12. A gym guide for machines and beginners
  13. Cross‑platform companions: web and iOS
  14. Practical engineering lessons and trade‑offs
  15. What FitGuru borrowed and what it built
  16. Where FitGuru goes next
  17. How FitGuru behaved in the wild: examples and edge cases
  18. Try it: code, APK, and demo
  19. Design checklist for teams building private on‑device coaches
  20. FAQ

Key Highlights

  • FitGuru splits the coach into two on‑device systems: a low‑latency Reflex Engine that analyzes live pose landmarks and a Cognitive Engine that generates post‑set debriefs from numeric summaries — ensuring camera frames never leave the phone.
  • The app runs MediaPipe Pose on CameraX, computes per‑joint angles and state machines to count reps and flag faults, applies a “Gym Shield” to avoid passersby interfering, and optionally calls Gemini and ElevenLabs for structured feedback and studio voices using only compact JSON payloads.
  • Local‑first storage (Room) preserves workout history and metrics offline; voice caching and TTS fallbacks maintain hands‑free guidance in poor network conditions.

Introduction

A lifter in the middle of a heavy set cannot scroll a screen or manually record every cue their trainer would give. Commercial solutions typically send video to cloud servers for analysis, which raises privacy concerns and introduces latency. FitGuru reframes the problem: can a phone act as a credible, real‑time coach while ensuring the camera feed and landmarks never leave the device?

FitGuru answers that question with a practical, engineering‑first approach. Built as a weekend hack for Innovhacks 4.0 by Team ARKA, the app demonstrates a viable architecture for private, multimodal workout coaching on commodity phones. It pairs lightweight, deterministic on‑device processing with optional cloud intelligence for richer post‑set feedback. The result is a coach that counts reps, calls out form faults in real time, accepts voice commands, and — when permitted by the user — produces a structured recovery and biomechanics plan through an LLM, all without sending raw video or images to the cloud.

The next sections unpack the design choices, the engineering trade‑offs, and the practical techniques that make FitGuru work in messy real‑world settings like crowded gyms and noisy training floors.

Why “two brains” beats “one big model”

The straightforward idea — stream camera frames to a large model and ask “is this a good squat?” — fails for three technical and practical reasons: latency, privacy, and device resource limits.

Latency: a frame arrives every ~33 ms. Any meaningful response must be near‑instant; offloading frames and waiting for a remote or heavyweight local LLM produces jittery or useless feedback mid‑set.

Privacy: sending video of a person exercising is a heavy ask. Many lifters will not permit continuous streaming of their bodies, faces, or gym surroundings.

Memory: trying to run a multi‑GB LLM alongside live camera pipelines causes native memory crashes on many phones.

FitGuru splits responsibility instead of aggregating it. The Reflex Engine runs on every frame and performs deterministic geometry: landmark ingestion, joint angle computation, state‑machine transitions, rep counting, and short spoken cues. The Cognitive Engine wakes after the set and consumes a compact numeric summary (about 4 KB of text) to produce a structured biomechanics report and recovery suggestions. That JSON — not video or landmarks — allows cloud models to give the kind of human‑readable narrative lifters expect while preserving privacy.

This separation preserves low latency for live guidance and gives the LLM the context it needs to produce high‑value advice without seeing a single pixel.

From camera to meaningful joint angles

FitGuru’s visual pipeline is CameraX → MediaPipe Pose Landmarker in LIVE_STREAM mode. The landmarker returns 33 normalized 3D landmarks per frame (nose, shoulders, elbows, wrists, hips, knees, ankles, toes, etc.). The app does not ask the model to classify movements; it computes geometry from the landmarks.

Angle computation is straightforward vector math: for a joint B formed by points A–B–C, the vectors BA and BC are computed and the interior angle at B is the arccosine of the dot product of the normalized vectors. This interior angle becomes the fundamental sensor.

For a squat the coach measures knee and hip interior angles. A deadlift uses hip hinge and knee angles. Push‑ups monitor elbow angles and the shoulder–hip–ankle alignment to detect sagging or piking. Running is reduced to alternating ankle positions and cadence. Each movement maps to a small set of primary angles and threshold rules.

Practical note: normalized pose coordinates depend on camera placement and subject anthropometry. FitGuru mitigates variation by requiring multiple complementary angle checks (e.g., both knee and hip thresholds must be satisfied for a valid squat “down”) and by demanding a clear stand‑up before a rep is counted.

Real‑world example: a taller lifter with a longer femur will naturally present different minimum knee angles at depth than a shorter lifter. FitGuru’s combined‑angle rules reduce false negatives and false positives by requiring consistent joint behavior rather than a single absolute threshold.

A rep is a state machine, not a lone frame

Counting repetitions from independent frames is fragile. A classifier that labels individual frames as “down” or “up” will double‑count, ignore lockouts, and mistake noise for movement. FitGuru treats each exercise as an automaton: a small deterministic finite‑state machine (FSM) tracks transitions between ready → down → up and emits a rep only on the rising edge.

Example squat state machine:

  • ready —(knee < 105° AND hip < 115°)→ down
  • down —(knee > 155° AND hip > 155°)→ up (+1 rep)

The same approach extends to deadlifts, pull‑ups, curls, and planks (plank scores are measured as good‑seconds ÷ total‑seconds rather than discrete reps). State machines simplify reasoning about partial reps, cheaty swings, and lockouts. They also provide stable inputs for spoken cues: a hip collapse in the hole triggers “keep your chest up,” while an excessive upper‑arm swing on a curl triggers “keep elbows pinned.”

Cues have a cooldown: FitGuru enforces a 2.5‑second cooldown between spoken corrections to avoid nagging. The reflex layer remains focused on concise, tactical guidance; it is not a running commentary.

Real‑world example: during a heavy set, small transient errors in joint detection occur. The FSM prevents a momentary mis‑detected “down” from incrementing the rep count until the full cycle completes, preventing phantom reps when someone briefly occludes the camera.

Gym Shield: locking onto the right person

Gyms are hostile to pose detection. Someone walks behind you; MediaPipe detects a second person and suddenly counts their movement. FitGuru implements a “Gym Shield” that ensures the system locks on to the primary lifter.

The process:

  1. Score each detected body by torso size and proximity to frame center.
  2. Anchor on the most prominent torso.
  3. If that anchor disappears (someone crosses the frame), freeze the current rep state for 45 frames (~1.5 s). Do not increment or switch athletes during this grace period.
  4. Only pick a new primary subject after the grace window expires.

Additionally, the HUD displays the Gym Shield state: e.g., “Gym Shield: Passerby Filtered • Holding Rep State.” This explicit feedback exists because initial demos lost sets to strangers before the filter was introduced.

Real‑world example: in a busy CrossFit box, the coach can freeze rep counting briefly when a spotter crosses behind the lifter; FitGuru holds the state rather than attributing those reps to an unintended person.

Local‑first data: privacy by design

FitGuru is local‑first by architecture and by policy. Video never leaves the device. Landmarks never leave the device. Workout history is persisted in a local Room database. The app’s only optional cloud interaction occurs after you finish a set and only if you supply a Gemini API key.

WorkoutSession is a local Room entity containing date, exercise type, reps, form score, calories, average heart rate, and optional weight. Weekly goals, daily challenges, the dashboard, and analytics all query local data. This ensures a usable offline experience: if the network dies mid‑session, the coach still counts reps and speaks cues.

Crucial constraint: when the app calls Gemini for a debrief, it sends a tightly constrained text payload — a small paragraph of numbers (exercise name, completed reps, form score, weight, primary fault, measured key angle). No frames. No landmark arrays. No faces. The payload is roughly 4 KB of text. The LLM returns a JSON structure with diagnosis, corrective cues, accessory drills and a recovery protocol. The cloud model behaves as a sports scientist reading a stat sheet rather than a camera operator.

When no Gemini key is configured, or when the request fails, the app still produces a rule‑based report: diagnosis, three cues, two mobility drills. The set is never held hostage by the network.

Real‑world privacy implication: lifters who refuse any cloud processing can use the entire real‑time coaching stack locally; they only forgo richer narrative debriefs and studio‑quality voice packs.

Post‑set cognition: Gemini 2.5 Flash and structured debriefs

FitGuru uses Gemini 2.5 Flash for the post‑set biomechanics report and recovery protocol, with 1.5 Flash as a fallback. The app requests strict JSON and strips Markdown or extraneous formatting if necessary.

Typical Gemini response schema:

  • diagnosis: concise root‑cause explanation
  • correctiveCues: array of 3 compact cues
  • accessoryDrills: 2–3 drills
  • recoveryWindowHours, strainAssessment, proteinGrams, hydrationMl, activeStretches: structured recovery guidance

This JSON feeds the “Hear Coach Review” bottom sheet. ElevenLabs then reads the diagnosis and the cues in the voice persona you selected.

The cloud LLM never receives visuals. It only receives numeric summaries and structured metrics derived by the on‑device reflex layer. That mirrors how an in‑person trainer might read a client’s set summary and offer targeted programming cues.

Real‑world example: after a set of 12 squats with a form accuracy score of 88% and a primary fault of “chest collapse,” Gemini might return a diagnosis focusing on thoracic mobility, corrective cues like “brace core, drive chest up,” and accessory drills such as banded pull‑aparts and thoracic extensions. Those become audible next‑action items rather than raw video commentary.

Giving the coach a voice: ElevenLabs, caching, and graceful fallback

Speech is the interface. Android’s built‑in TTS can announce a rep, but it lacks nuance. FitGuru integrates ElevenLabs (eleven_turbo_v2_5) to produce studio‑grade voices with three personas: Coach Marcus (drill sergeant), Coach Maya (mindful guide), and Coach Alex (Olympian).

Latency and reliability are critical:

  • The voice stack is cache‑first. Each phrase is hashed (voiceId + phrase → MD5). If an MP3 exists in the cache, Play that immediately (0 ms path).
  • On cache miss, POST the text to ElevenLabs, store the resulting MP3, and then play it. This first‑time request can take several seconds (standard connect/read timeouts: 8 s / 12 s). Repeated, frequent phrases become instant.
  • On timeout, HTTP error, or no API key, fall back to Android TextToSpeech.

Download‑then‑play (rather than streaming) avoids jitter under variable network conditions. That yields a consistent experience: during a heavy set, a cached “Good lockout!” plays instantly; first‑time longer debriefs may take a moment but are still accessible. Offline, the voice becomes flatter, but never silent.

Guardrails prevent self‑triggering: when the app is speaking, the speech recognizer ignores the mic to avoid the coach “hearing” itself and executing a voice command.

Real‑world example: a lifter who prefers a motivational voice for top‑set cues selects Coach Marcus; their common mid‑set cues are cached over several workouts and play with imperceptible latency, preserving focus.

Hands‑free controls tuned for the gym

FitGuru’s VoiceCommandManager wraps Android SpeechRecognizer and keeps listening in a restart loop. The set of hands‑free commands emphasizes operational needs during a session:

  • “finish set” / “end workout” → stops camera loop, opens Gemini summary
  • “smart rest” / “take a break” → opens rest timer and reads heart rate
  • “form check” / “coach help” → speaks a current lift cue
  • “switch to squat” → changes the state machine without leaving the HUD
  • “weight 40” → sets the load for volume math

Two guardrails harden robustness:

  1. Echo lock: ignore the mic while ElevenLabs or TTS is speaking to prevent the coach from hearing its own voice.
  2. 2‑second command debounce: prevent partial recognition from firing one command multiple times.

Real‑world example: while holding a barbell in a heavy rack position, the lifter can say “finish set” without dropping the bar or touching the phone; FitGuru closes the camera loop and prepares the post‑set debrief for the user.

Heart rate as a second sense

Vision provides form; heart rate provides physiological state. A Wear OS / Galaxy Watch client pushes BPM to the phone over the Wearable Data Layer path (/heart_rate). The DataLayerListenerService broadcasts that into the live HUD.

Smart Rest uses heart rate in a deliberately simple, actionable way:

  • If heart rate is at or below 120 BPM, FitGuru says you are ready to lift.
  • If above 120 BPM, FitGuru recommends breathing and offers an extra 30 seconds.
  • When the timer reaches zero, the coach announces “Rest complete.”

Heart rate also flows into the WorkoutSession persisted to Room and rides with the session into the Gemini recovery prompt. The system is optional; the camera coach does not depend on it.

Real‑world example: during a cluster set, a lifter’s form score drops while BPM climbs. Future plans include real‑time heart‑rate fusion so the coach can call out tactical adjustments mid‑set (e.g., “form slipping—reduce weight or rest longer”).

A gym guide for machines and beginners

A first‑time lifter needs more than rep counting. FitGuru bundles a local gym_catalog.json with 20 common machines and movements (lat pulldown, seated row, chest press, shoulder press, leg press/extension/curl, Smith, squat rack, dumbbells, bench, pec deck, treadmill, bike, rower, mat, deadlift platform, pull‑up bar, running cadence).

Each card includes:

  • what the machine/movement is
  • primary muscles
  • setup steps
  • execution steps
  • common mistakes
  • a camera tip

If a movement has an associated state machine, tapping the card launches live coaching. The home screen surfaces weekly goals, a rotating daily challenge, quick analytics (form/volume/HR), and shortcuts into the gym directory. Profile and history remain local; no cloud account is required.

Real‑world example: a new gym member can select “lat pulldown,” follow setup cues to adjust the seat and strap, then enable live coaching for reps and form reminders — all without trainer supervision.

Cross‑platform companions: web and iOS

FitGuru’s judges‑friendly web companion is a Node/Express app with MediaPipe in the browser, a skeleton overlay, the same cue dictionary, and an optional wearable BPM simulator. It deploys as a Docker container. iOS is a SwiftUI client that leverages Apple Vision (VNDetectHumanBodyPoseRequest) with the same cloud integrations (ElevenLabs, Gemini), gym catalog bundled, and local history in on‑device storage.

The core idea remains identical across platforms: pose on device, language after the set, voice as the interface. Android contains the deepest implementation because CameraX and MediaPipe wiring is stable there, but the approach generalizes.

Real‑world example: a developer jury can run the web companion locally in a browser to demo the MediaPipe pose loop without needing Android Studio.

Practical engineering lessons and trade‑offs

FitGuru was built under a weekend hackathon constraint with a single inflexible rule: video must never leave the phone. That constraint shaped the architecture and surfaced valuable lessons.

  1. Live camera + large local LLM do not coexist comfortably. Running a ~1.3 GB model while CameraX + MediaPipe was active caused native memory crashes. FitGuru added pause()/resume() around the landmarker so heavy inferences can run in isolation. The mid‑set “Gemma AI” is a templated cognitive layer; the real LLM runs after sets.
  2. Wearable integrations must use stable APIs. An initial attempt with Samsung’s Sensor SDK failed; switching to the Wearable Data Layer (/heart_rate) simplified reliability and cross‑device compatibility.
  3. Fixed thresholds fail across bodies. A knee < 105° rule works for many but not all. FitGuru tightened rules by requiring complementary angles and full stand‑ups before counting, and by adding a 2.5 s cue cooldown. Per‑user calibration remains the honest next step.
  4. Voice + microphone introduces feedback loops. The coach initially “heard” itself and executed voice commands. Echo lock and debounce fixed that.
  5. Crowded gyms stole reps until Gym Shield was implemented. Prominence scoring and a 45‑frame occlusion hold turned demos into resilient gym behavior.

These lessons reflect real constraints when moving from clean demos to messy deployments: memory is finite, sensors are noisy, and human environments are unpredictable.

What FitGuru borrowed and what it built

The team began from Google’s MediaPipe Pose Landmarker Android sample and retained CameraX + landmarker wiring. That saved substantial low‑level work around YUV buffers and camera plumbing.

Everything above that sample is original:

  • exercise state machines
  • Gym Shield
  • Room history and goals
  • gym catalog
  • Gemini debrief integration
  • ElevenLabs voice caching and personas
  • VoiceCommandManager and Smart Rest
  • analytics, web companion, and iOS client

Acknowledging a robust sample is good engineering discipline; differentiating product from example code is honest design.

Where FitGuru goes next

FitGuru’s roadmap focuses on making the coach smarter and more personalized while retaining privacy guarantees.

Planned improvements:

  • On‑device LLM: a smaller, efficient on‑device model (Gemma or MediaPipe LLM Inference) using the pause/resume path for offline debriefs without an API key.
  • Live heart‑rate fusion: integrate BPM into mid‑set guidance rather than limiting it to rest timers and recovery prompts.
  • Per‑user angle calibration: film a single clean set and tailor thresholds to that skeleton to reduce false positives/negatives across body types.
  • Session time‑series: richer continuous stores for form scores vs heart rate over minutes, enabling deeper analytics and trend detection.

These steps aim to increase accuracy and personalization while preserving the core privacy posture.

How FitGuru behaved in the wild: examples and edge cases

Crowded CrossFit box: initial demos often credited passersby with reps. Gym Shield and occlusion holding made the system reliable enough for real sessions.

Outdoor runs: camera‑based cadence coaching works but depends on phone placement. FitGuru recommends chest or waist placement for stable ankle visibility; cadence alerts trigger under 155 steps per minute.

Home workouts with children: when a child briefly crosses the scene, the anchor remains until the 45‑frame timeout expires; the system avoids switching primary athletes mid‑set.

Heavy compound sets: for deadlifts and squats, the state machine avoids half‑rep counting and requires clear lockouts before incrementing reps. Spoken cues reduce the need to peek at summaries between sets.

Network outages: voice still functions via Android TTS fallback; Gemini debriefs are disabled when no API key or network exists, but a rule‑based local report ensures actionable feedback always appears.

Try it: code, APK, and demo

Team ARKA published the project on GitHub and provided an APK for quick testing. The repo contains instructions for building, the web companion, and deployment specs.

Resources:

Optional settings:

  • Gemini API key for post‑set debriefs
  • ElevenLabs API key for studio voices

Neither key is required for the core reflex coach and rep counting.

Design checklist for teams building private on‑device coaches

If you plan to build a similar product, FitGuru’s pragmatic design checklist helps navigate trade‑offs:

  1. Separate real‑time sensing from higher‑order reasoning. Keep a deterministic, low‑latency layer for live guidance and an optional cognitive layer for post‑set analysis.
  2. Use geometric, interpretable features. Joint angles and state machines are robust and explainable compared to opaque frame classifiers.
  3. Implement a prominence anchor and occlusion grace period to handle multiple people and transient occlusions.
  4. Make cloud calls opt‑in and constrained. Send only small numeric summaries for LLM consumption, never raw video or landmarks.
  5. Cache voice clips aggressively. Prioritize cached audio playback for frequently used phrases.
  6. Design for degraded networks. Provide local rule‑based fallbacks for every cloud feature.
  7. Plan for per‑user calibration. Anthropometry varies; per‑user adjustments improve accuracy.
  8. Test in realistic environments: noisy gyms, crowded spaces, and with wearables paired.

FAQ

Q: Does FitGuru ever upload my camera video or pose landmarks? A: No. Video frames and pose landmarks remain on the device. Only an optional post‑set JSON summary (text, ~4 KB) is sent to Gemini if you supply an API key.

Q: What devices are supported? A: The Android APK targets API 29+. FitGuru relies on CameraX and MediaPipe Pose for Android. There is also a web companion and an iOS SwiftUI client (using Apple Vision) with comparable features.

Q: How accurate is the rep counting? A: Accuracy depends on camera placement, lighting, and user anthropometry. FitGuru’s FSMs and combined‑angle thresholds reduce false counts versus single‑frame classifiers, but edge cases exist. Per‑user calibration is planned to improve consistency.

Q: What happens without an internet connection? A: The reflex coach (pose processing, rep counting, local cues, and local history) works offline. Speech falls back to Android TextToSpeech, and Gemini debriefs are unavailable; however, a local rule‑based report still appears.

Q: Can I use my smartwatch? A: Yes. A Wear OS or compatible watch can push heart rate over the Wearable Data Layer to the phone. Heart rate informs Smart Rest and is included in session summaries.

Q: Do I need Gemini or ElevenLabs keys? A: No. Both are optional. Gemini provides richer post‑set narratives and recovery advice; ElevenLabs provides studio‑quality voices. If neither is configured, the app still counts reps, gives live cues, and uses UI‑based or TTS guidance.

Q: How does FitGuru handle multiple people in frame? A: FitGuru scores detected bodies by torso size and centrality, anchors on the most prominent torso, and holds an occluded state for 45 frames (~1.5 s) to avoid switching athletes during crosses.

Q: How customizable are the cues and thresholds? A: At present, thresholds are fixed per movement with conservative combined checks. The team plans per‑user calibration (film a clean set, fit thresholds) and additional personalization on the roadmap.

Q: Is the app open source? A: Yes. The source code is available: https://github.com/atharvachaudhari1/innovhacks4.0_arka. The repo includes the Android project, web companion, and instructions for building and testing.

Q: How does FitGuru avoid sounding repetitive during sets? A: Spoken cues have a 2.5‑second cooldown, voice responses are short, and the voice cache ensures immediate playback of common phrases without network latency. The app avoids praise on every rep to prevent distraction.

Q: Can I use FitGuru for programming or progressive overload tracking? A: FitGuru stores reps, volume (weight × reps), form score, and average heart rate locally. These metrics feed weekly goals, daily challenges, and local analytics. Advanced programming and cloud sync are not part of the initial design.

Q: What are the main limitations? A: Limitations include per‑user anthropometric variance, device memory constraints (preventing large on‑device LLMs from running concurrently with the camera), and dependence on camera angle for certain lifts. The team has identified paths to address each limitation.

Q: How can I help improve or contribute? A: The code is on GitHub. Contributions, issues, and PRs are welcome for state machines, calibration routines, voice cache strategies, and additional exercise support.


FitGuru demonstrates a pragmatic path toward private, on‑device movement coaching. Deterministic geometry, small state machines, and careful guardrails deliver reliable real‑time guidance. Optional cloud cognition enriches the post‑set experience without compromising privacy. This combination makes it possible to have a pocket coach that counts reps, corrects form, and keeps the video in your hands.

RELATED ARTICLES