How Body Forge Built an AI Workout Coach: One‑Tap Logging, GPT‑4o Analysis, and a Minimalist In‑Session UX

Building an AI workout coach in SwiftUI: one-tap logging during the set, GPT-4o analysis after

Table of Contents

  1. Key Highlights:
  2. Introduction
  3. Designing the in‑workout screen for the worst moment
  4. One‑tap logging: compliance, context, and data quality
  5. From noisy logs to structured summaries: what to send to the model
  6. Prompting GPT‑4o for actionable coaching (not fluff)
  7. The stack under the hood: SwiftUI, Core Data, CloudKit, HealthKit
  8. Privacy, security, and ethics of sending workout data to an LLM
  9. Programme library, exercise content, and educational assets
  10. Landing page strategy: convert, clarify, and load fast
  11. Lessons learned and engineering trade‑offs
  12. Roadmap considerations and future directions
  13. FAQ

Key Highlights:

  • Body Forge prioritizes in-session simplicity: one‑tap logging and a minimal UI preserve set-by-set context without distracting the user.
  • Post-session analysis is driven by compact, structured summaries sent to GPT‑4o, producing fast, concrete coaching suggestions while controlling token costs and privacy exposure.
  • The native iOS stack (SwiftUI, Core Data + CloudKit, HealthKit) supports offline-first reliability, cross‑device history, and heart‑rate integration, with a focused landing page to convert search traffic into installs.

Introduction

The gap between raw gym logs and useful coaching lies in context. Most people capture numbers in generic notes or trackers that record sets and repetitions but forget the narrative: how weights change over weeks, when progress stalls, or when programming adjustments are necessary. Body Forge approaches that problem with a clear rule: don’t ask users to think during the set; analyze after it.

That rule reorganizes the app experience around two distinct moments. During the set, every element on screen must survive a sweaty thumb and ten seconds of attention. After the session, a compact, structured summary feeds a large language model (GPT‑4o) that can compare recent trends and deliver targeted, actionable advice. The result: a coach that fits in a pocket and reacts to real, measurable progress.

The following piece examines the design choices, data model, backend constraints, and product decisions behind that approach. It translates them into practical guidance for teams building similar fitness tools, and explains why a modest amount of preprocessing changes the quality of AI coaching dramatically.

Designing the in‑workout screen for the worst moment

Exercise sessions are noisy. Hands get slippery, minds wander between sets, and motivation peaks and troughs on every rep. The critical design decision is to make the UI resilient to that environment. Body Forge reduces cognitive load by showing almost nothing while a set is happening. That minimalism is neither aesthetic nor ideological; it solves real retention and accuracy problems.

What a user sees during a set:

  • Last week's sets for this exercise — immediate context that informs loading and intensity.
  • A rest timer — crucial for pacing between sets.
  • Weight and reps fields — easily editable without switching screens.
  • One tap to log the set — a single primary action that ends the interaction.

Why that matters

  • One‑handed use: The app must be fully operable with a thumb. Big, clear tappable areas, large numbers, and minimum navigation prevent accidental taps and frustrated users.
  • Attention cost: After finishing a set, most people want to rest; they rarely want to parse charts, tweak advanced settings, or read paragraphs. The design respects that limited attention and defers analysis until after the workout.
  • Data fidelity: Simple, fast logging increases the likelihood that entries happen at the right time and with correct values.

Real‑world comparisons

  • Many popular trackers place charts, complex navigation, and detailed analytics right in the session screen (e.g., some All‑In‑One fitness apps). This adds friction. Users either delay logging or move to notes apps that lack temporal continuity.
  • Apps like Strong or Jefit offer quick logging but often encourage richer in‑session tweaks, which can slow the flow. Body Forge’s decision is stricter: keep the session screen minimal and push analysis to the quiet window after the workout.

Design patterns to replicate

  • Primary action emphasis: One dominant “Log” button visible with a high‑contrast color.
  • Progressive disclosure: Expose only necessary controls; hide advanced options behind gestures or secondary screens accessed between workouts.
  • Immediate context: Show previous set(s) of the same exercise so users can make informed next‑set decisions without diving into history pages.

One‑tap logging: compliance, context, and data quality

Logging behavior shapes the data your model sees. If users delay or approximate their entries, downstream analysis becomes noisy. One‑tap logging reduces friction and yields higher‑quality inputs.

Behavioral benefits

  • Reduced abandonment: When logging is fast, users are more likely to record every set, preserving the sequence and nuances that matter for trend analysis.
  • Time‑consistent entries: Immediate logging timestamps sets correctly, enabling accurate rest interval calculations and tempo analysis when applicable.
  • Lower editing overhead: Users make fewer post‑session corrections, reducing inconsistencies across sessions.

Implementation details

  • Large tappable targets mapped to both screen and accessory inputs (e.g., Apple Watch, HomePod shortcuts) expand options for hands‑occupied scenarios.
  • Automatic defaults: Pre-fill the next set’s weight based on a simple progression rule (e.g., +2.5 kg when previous set achieved target reps) while allowing quick overrides.
  • Haptic feedback: Subtle confirmation reduces cognitive load by signaling success without visual attention.

Potential pitfalls

  • Accidental logs: Use undo affordances or brief “logged” toasts with a 3‑second undo window to prevent permanent errors from slips.
  • Overautomation: Don’t auto-adjust future loads aggressively. Provide conservative defaults and let the model recommend larger programming changes.

Real example: how small UX details affect data A gym user who has to pinch to open a menu to log a set will often delay logging until after training. Those delayed entries cluster and lose the per‑set context necessary to detect trends like accumulated fatigue or a failure to hit a target rep range. One‑tap logging reduces that latency and preserves the temporal structure of the workout.

From noisy logs to structured summaries: what to send to the model

Large language models excel when they receive clean, structured data. Raw event logs from sessions—complete with timestamps, empty fields, and manual edits—are noisy and expensive to process. Body Forge creates compact WorkoutSummary objects that distill what matters: volume, intensity, progression over recent weeks, and per‑exercise details.

A concise data model The app transforms raw logs into a compact, typed summary before calling GPT‑4o. A simplified schema looks like:

struct ExerciseSummary: Codable { let name: String let sets: [SetEntry] // weight, reps, rest let topSetLast3Weeks: [Double] let volumeChange: Double // vs. previous session } struct WorkoutSummary: Codable { let date: Date let programme: String // e.g., "Upper/Lower, week 4" let exercises: [ExerciseSummary] let avgHeartRate: Int? }

Why structure matters

  • Cost control: Compact numerical arrays consume far fewer tokens than verbose logs, reducing cost and improving response time.
  • Model reasoning: Numeric arrays and labelled fields let the model compute trends (e.g., linear progression, plateau detection) reliably instead of guessing intent from free text.
  • Reproducibility: Structured inputs enable deterministic heuristics and hybrid workflows—some decisions can be made by deterministic code and only ambiguous cases passed to the model.

Key metrics to include

  • Volume (sets × reps × load) per exercise and total.
  • Intensity: relative percentage of maximum or prescribed loads.
  • Top sets across recent weeks for progression detection.
  • Rest intervals when relevant (for hypertrophy or strength programming).
  • Heart rate data when available (from HealthKit) to correlate effort with physiological cost.

Preprocessing techniques

  • Outlier removal: Automatically discard improbable values (e.g., 1000 kg entry) before summarization.
  • Smoothing: If users perform minor deloads or spikes (travel, illness), present the model with a simple moving average and list exceptions separately.
  • Context flags: Mark sets performed as RPE-reached failures, warmups, or forced reps to help the model interpret the numbers.

Practical example Consider an exercise where the top set weights for the last three weeks are [110, 112.5, 112.5] and the current session top set is 110. Volume change is −3% vs. the previous session. That pattern signals a likely stall. Sending the model this compact summary allows it to recommend concrete programming: a planned deload, technique check, or a micro‑cycle to restore progression.

Prompting GPT‑4o for actionable coaching (not fluff)

Not all AI feedback is useful. Vague encouragement has no place in coaching; users need a decision pathway. The prompt construction behind the app’s communication with GPT‑4o must be deliberate: precise instructions, constrained output formats, and an emphasis on concrete next steps.

Prompt design principles

  • Ask for short, specific recommendations. Request 2–3 prioritized actions (e.g., “deload to 80% next week,” “reduce weekly volume by 10%,” “switch to 6–8 rep range”).
  • Provide the numerical evidence to justify recommendations. Let the model reference top sets, volume changes, and heart rate trends.
  • Constrain the model’s output to a predictable structure (e.g., JSON with keys like “summary,” “recommendation,” “confidence”).
  • Use a hybrid approach: let deterministic code make obvious decisions (e.g., when volume increases by >20% in a session), and ask the model to handle ambiguous cases.

Example prompt (conceptual) You might send GPT‑4o a task: “Analyze the following WorkoutSummary. For each exercise, determine whether to (A) continue progressive overload, (B) deload, (C) reduce volume, or (D) change rep range. Provide a short rationale and a concrete action for next session (weights or percentages). Base decisions on topSetLast3Weeks and volumeChange. Output JSON.”

Benefits of constrained outputs

  • Predictable UI rendering: The app can display the model’s advice in a consistent card format without parsing arbitrary text.
  • Easier testing: Generate synthetic WorkoutSummary objects and validate that the model’s outputs match expectations.
  • Safer guidance: The model is less likely to hallucinate training jargon if constrained to produce specific labels and numeric prescriptions.

Balancing speed and quality

  • Latency target: Body Forge aims for the response around 10 seconds. That is acceptable for post‑session reflection. If responses take longer, provide an intermediate “processing” UI and show cached heuristics in the meantime.
  • Cost tradeoffs: Limit the number of exercises per call or summarize only the most critical ones if you want lower token usage for free users.

Examples of model outputs that matter

  • Concrete: “Incline press stalled for three weeks. Deload to 80% of last top set for two sessions, then attempt a 2.5 kg increase on session 3.”
  • Unhelpful: “Keep pushing hard and try harder next time.” Avoid outputs like this by forcing specificity and numbers.

The stack under the hood: SwiftUI, Core Data, CloudKit, HealthKit

Choosing the native iOS stack reflects a decision to optimize for fluid UI performance, battery life, and tight integration with Apple services. Body Forge’s stack supports offline-first usage, multi-device sync, and physiological data capture.

SwiftUI for responsive, single‑hand interactions

  • Declarative UI: SwiftUI simplifies state-driven interfaces, enabling quick iteration on minimalist designs and animations that cue users without verbosity.
  • Platform ergonomics: SwiftUI supports accessibility out of the box (Dynamic Type, VoiceOver), which is essential when users interact with the app with limited attention.

Core Data with CloudKit for history and sync

  • Local persistence: Core Data stores session logs locally for immediate access even in poor connectivity scenarios.
  • CloudKit sync: Pairing Core Data with CloudKit ensures users retain their exercise history when switching devices or restoring from backups. Design the sync model carefully to avoid conflicts on concurrent edits.
  • Data modeling best practices: Keep the schema stable, use lightweight migrations, and avoid deeply nested relationships that complicate conflict resolution.

HealthKit integration for heart rate

  • Heart rate correlation: Average heart rate during a session adds a physiological signal to training load and perceived exertion. It helps detect overreaching or under‑effort.
  • Privacy and permissions: Request only necessary permissions. Clearly explain why heart rate is used and make its reporting optional.
  • Sampling and smoothing: Raw heart rate data is noisy. Report average HR, peak HR, and HR variability markers if available, but keep the primary input simple (avgHeartRate) to limit complexity.

Infrastructure considerations

  • Local-first analysis: Do as much preprocessing and deterministic checks on device as possible to reduce cloud calls and latency.
  • Model calls: Send compact WorkoutSummary objects to GPT‑4o endpoints. Avoid sending identifiable data or long chat histories.
  • Caching: Store the model’s last response per workout so users can revisit advice offline.

Resilience patterns

  • Retry strategies for model API failures with exponential backoff.
  • Fallback heuristics if the model is unavailable: simple rule-based messages such as “increase by 2.5 kg if you hit target reps” maintain the product promise.
  • Data integrity checks: Validate set entries before saving and provide an undo stack for accidental logs.

Privacy, security, and ethics of sending workout data to an LLM

Sending user data to an external model introduces privacy and ethical responsibilities. Small changes in the data you transmit protect users and preserve model usefulness.

Minimize transmitted data

  • Send only aggregated, non‑personalized numbers. Avoid names, location, raw timestamps beyond session date, or PII.
  • Exclude biometric identifiers. Average heart rate is useful; raw heart rate time series is usually unnecessary.

Consent and transparency

  • Communicate what is sent: Provide a short privacy prompt at first model use explaining that a compact workout summary will be sent to generate advice.
  • Allow opt‑out: For users who prefer local-only functionality, offer deterministic heuristics as an alternative.

Secure transmission and storage

  • Encrypt data in transit (HTTPS/TLS) and enforce strict server endpoint policies.
  • Do not store model responses long term unless necessary. If stored, encrypt and give users control over deletion.

Regulatory considerations

  • Health data may fall under regulatory scrutiny in some jurisdictions. Treat workout and heart rate data with the same care as other sensitive personal data and consult legal counsel when scaling internationally.

Ethical coaching

  • Avoid prescriptive medical claims. The model should not diagnose or attempt to treat injuries.
  • Flag edge cases: If the model detects anomalous patterns that could indicate risk (e.g., acute HR spikes), present a cautious message recommending professional assessment rather than definitive statements.

Programme library, exercise content, and educational assets

Accurate coaching blends data with programming structures. Body Forge uses a programme library and an exercise library to situate each workout within a broader training plan.

Programme models

  • Template programmes: Upper/Lower splits, full‑body beginner cycles, and block‑based progressions cover common training needs.
  • Programmatic week tracking: A structured “programme” field in the summary helps the model interpret whether a session is meant to be an intensity week, a deload, or an accumulation phase.
  • Custom blocks: Allow power users to create and save custom blocks (e.g., hypertrophy block, strength block) that the model can reference.

Exercise library

  • Catalog: Over 200 exercises with video demos teach proper form and reduce injury risk.
  • Linking the summary: Each ExerciseSummary includes a canonical exercise name or ID so the model can apply exercise‑specific heuristics (e.g., bench press vs. leg press programming differs).
  • Localized content: Videos and cues available in multiple languages increase accessibility.

Educational overlays

  • Short form content: Use small inline tips that explain why a recommended deload is appropriate or what “80%” means in practice.
  • Modal deep dives: Provide optional expanded content for users who want a deeper understanding of periodization or autoregulation.

Real example: When a programme and data point intersect If a user is on “Upper/Lower, week 4” and the model sees a three‑week stall on incline press, recommendations should factor in the phase of the programme. During a planned intensity peak, the model might advise switching to an accumulated micro‑cycle rather than immediate deload.

Landing page strategy: convert, clarify, and load fast

An App Store link alone rarely conveys product differentiation. Body Forge launched a dedicated bilingual landing site with a single conversion goal per screen: go to the App Store. The site’s design decisions support quick discovery and trust.

Single goal per screen

  • Focus reduces friction. Each scroll frame answers one question: what the app does, how it works, trust signals (ratings), and a final call to action.
  • Bilingual approach reaches more users without complicating the narrative.

Performance matters

  • The first screen paints in about 0.7 seconds. Fast perceived load time improves both retention and SEO.
  • Optimize hero images, use a minimal CSS baseline, and preconnect to the App Store where necessary.

SEO and App Store discoverability

  • Landing pages can capture search queries that the App Store cannot easily target (e.g., “AI workout coach one‑tap logging”).
  • Structured data and clear headings help search engines index the product’s unique value proposition.

Conversion copy

  • Avoid vague marketing claims. Use concrete examples of the app’s benefits: “One‑tap logging, post‑session AI coaching in ~10 seconds, CloudKit sync.”
  • Show example outputs: a short coaching card with a concrete recommendation increases credibility.

Tracking and privacy

  • Keep analytics minimal. Use privacy‑respecting analytics to measure conversion without over‑tracking prospects.

Lessons learned and engineering trade‑offs

Body Forge’s approach yields several transferable lessons for product teams building AI‑assisted fitness tools.

Design the in‑session experience for the worst moment

  • Expect low attention and single‑hand interaction.
  • Prioritize speed and low friction over features that only power users appreciate.

Summarize before you prompt

  • Structured, numeric summaries enable precise model reasoning, reduce token usage, and allow hybrid rule‑based decisions.
  • Avoid sending raw logs unless necessary for rare, complex cases.

Make advice concrete

  • Users respond to specific, numeric direction. “Deload to 80% next week” is actionable; generic motivational sentences are not.
  • Enforce output constraints on the model to avoid ambiguity and keep UI rendering predictable.

Invest in offline reliability and sync

  • Local-first storage prevents data loss and preserves the user’s history across devices.
  • Use CloudKit or equivalent to provide seamless accountless sync on Apple devices.

Balance automation and control

  • Provide thoughtful defaults but leave control to the user. Automation should reduce effort, not erase agency.

Cost and latency management

  • Model costs can scale with user base. Use preprocessing to narrow the model’s input and reduce token consumption.
  • Limit model calls—only run analyses at the end of sessions or when a threshold of new data justifies a reanalysis.

Examples from other products

  • Fitbod and Strong use robust exercise libraries and strong in-session UX, but typically rely on deterministic progression models rather than LLM analysis. Body Forge demonstrates a hybrid path: a simple in-session UX plus an AI layer that provides contextualized advice.
  • Some health platforms push too much in-session information, leading to reduced adherence. Minimalism improves logging rates and quality of historical data.

Roadmap considerations and future directions

Scaling an AI workout coach involves technical, UX, and ethical evolution. Possible extensions that preserve the core philosophies include:

Personalized model tuning

  • Fine-tune small models on anonymized, consented user data to specialize coaching for populations (e.g., natural lifters vs. powerlifters).
  • Maintain strict opt-in and clear benefit explanations.

Edge inference

  • Investigate on-device LLMs for immediate offline analysis and privacy-preserving coaching for basic recommendations, reserving cloud LLMs for advanced analysis.

Richer physiological signals

  • Incorporate HRV, sleep, and recovery metrics if users opt in. Use them to contextualize performance dips and suggest recovery strategies.

Coach‑user conversations

  • Add a short Q&A flow where users can ask clarifying questions: “Why did you recommend a deload?” The model can be constrained to cite the numeric evidence provided in the WorkoutSummary.

Social and accountability features

  • Group programmes, shared progression charts, or coach‑moderated plans can improve adherence. Keep these optional to avoid feature bloat.

A/B testing of guidance styles

  • Test different tones and recommendation granularities to find the mix that drives adherence and perceived usefulness. Concrete, short directives typically outperform long narratives.

FAQ

Q: How does one‑tap logging handle warm‑up sets and assistance work? A: The app distinguishes warm‑ups and working sets either through user tagging or automatic heuristics (e.g., sets significantly below recent top sets or marked with short rest intervals). Only working sets typically feed the WorkoutSummary metrics used for progression and volume calculations; warm‑ups can be recorded but are flagged separately.

Q: Is my workout data stored on servers or kept on my phone? A: Core session logs are stored locally using Core Data. If you enable sync, CloudKit replicates data across your Apple devices. Model requests send only a compact, non‑identifiable WorkoutSummary to the external model endpoint; raw logs and media (e.g., video) are not transmitted by default.

Q: Why send a summary to GPT‑4o instead of using on‑device heuristics? A: Deterministic heuristics are suitable for many simple rules (e.g., automatic microload increases). LLM analysis provides contextually rich recommendations when trends are ambiguous: it can weigh multiple signals simultaneously (volume, intensity, programme phase, heart rate) and suggest prioritized actions in natural language. The hybrid approach uses heuristics for clear cases and the model for nuanced ones.

Q: How long does it take to get feedback after a session? A: Typical response time is around 10 seconds. This latency balances cost, model complexity, and user expectations for post‑session reflection. If the model is unavailable, the app displays deterministic fallback recommendations.

Q: How are model costs managed as the user base grows? A: Cost control comes from two tactics: pre‑summarizing logs into compact numeric objects and limiting the number of model calls (e.g., one post‑session call rather than per‑exercise). Additional strategies include tiered features (premium users receive deeper analysis) and batching multiple workouts for less frequent, more comprehensive reviews.

Q: How does the app avoid giving unsafe advice? A: The app’s LLM prompts include safety constraints and disclaimers. It avoids medical diagnoses and recommends professional consultation when encountering patterns outside benign training fluctuations (e.g., persistent heart rate anomalies). The model’s outputs are also validated client‑side for format and content safety before display.

Q: How are programmes and exercises localized? A: The app ships with Russian and English localizations for UI and exercise content. The exercise library includes localized video demos and cues. When scaling to more languages, prioritize program names, exercise nomenclature, and measurement units (kg vs. lb) to maintain clarity in coaching recommendations.

Q: Can custom programmes be used with the LLM analysis? A: Yes. Custom blocks are included in the WorkoutSummary’s programme field, allowing the model to interpret sessions within the user’s chosen structure. Users can create templates that the model references when generating advice.

Q: What happens if I miss sessions or travel? A: The preprocessing step smooths short absences and marks exceptions. When the model sees abrupt changes in volume or intensity due to missed sessions or travel, it can recommend a structured reintroduction (e.g., reduced volume week followed by progressive increases) rather than assuming failure.

Q: How does heart rate data improve recommendations? A: Average heart rate during sessions gives a proxy for physiological load. Elevated HR at constant weights may signal increased fatigue or recovery deficits; low HR with declining performance might indicate insufficient effort or external stressors. The model uses avgHeartRate as an optional contextual feature to adjust recommendations.

Q: Why a dedicated landing page instead of relying on the App Store description? A: The App Store listing cannot explain nuanced differences or capture search intent effectively. A focused landing page clarifies unique value propositions (one‑tap logging, AI analysis) and improves SEO. Fast load times and a single conversion goal per screen increase install conversion rates.

Q: Are there plans for Android or cross‑platform support? A: The initial design and stack focused on iOS to leverage SwiftUI, HealthKit, and CloudKit for a native experience. Cross‑platform expansion requires rethinking sync, health data sources, and UI paradigms. Possible paths include native Android development with Google Fit integration or a backend‑centric approach with platform‑agnostic clients.

Q: Can I share model responses with a human coach? A: Exporting a session summary or the model’s recommendations as a PDF or shareable link is a common feature. It lets users present structured, evidence‑based summaries to a coach for deeper consultation without sending raw logs.

Q: How are recommendations updated as I progress? A: Every new session produces a new WorkoutSummary reflecting the latest three‑week context. The model takes the rolling window into account, so recommendations evolve as trends emerge—promoting progressive overload when appropriate and prescribing deloads or volume adjustments when necessary.

Q: What if I disagree with a recommendation? A: Users can override advice and log their chosen weights and reps. The app treats the model’s output as guidance, not a mandate. A feedback mechanism that records whether users follow recommendations helps refine defaults and informs future model prompts.

Q: How does the product handle conflicting signals across exercises? A: The model is instructed to prioritize systemic recommendations when multiple exercises show similar trends (for example, a high total session volume with HR elevation might trigger a recommendation to reduce overall volume rather than tweak a single exercise). Per‑exercise actions follow after systemic decisions.

Q: Are there analytics for coaches or remote trainers? A: Core product focuses on individual users. However, a coach dashboard is a plausible extension, where consented users share structured summaries with a coach. That feature requires additional privacy controls, audit logging, and explicit consent flows.

Q: Can the app help with technique or injury risk? A: The current approach relies on numeric summaries and library content. Video analysis or live form coaching would require additional tooling (computer vision, longer videos) and carries higher privacy and safety implications. Those are natural future directions but demand careful design and explicit opt‑in.

Q: How does the app recommend percentages like “80%” without knowing 1RM precisely? A: Percentages are typically calculated relative to recent top sets or programmed targets rather than estimated 1RM. If the programme includes a defined max or 1RM estimate, the model uses that; otherwise, it bases percentage suggestions on recent top set weights for safety and practicality.

Q: What are the next steps for someone building a similar app? A: Start by minimizing in‑session friction and collecting high‑quality local data. Create a compact summary model and prototype a few prompt templates. Iterate on constrained model outputs and validate recommendations with small user cohorts. Keep privacy at the center of every decision.


This detailed overview explains why a simple rule—keep the session screen minimal and analyze later—reshapes both the user experience and the technical architecture. It shows how structured data, constrained prompts, and native iOS capabilities combine to deliver fast, actionable coaching that users can actually use.

RELATED ARTICLES