How Zone Pedal
Works
How selected lower-zone steady work can use HR feedback, why short intervals stay power-led, and what the current control and learning evidence does — and does not — prove.
In plain English
On selected lower-zone steady phases, Zone Pedal can guide compatible smart-trainer resistance using heart rate as the feedback signal. That control requires usable resistance headroom above the trainer's floor, and the current envelope still has open arrival failures. Short intervals stay power-led because heart rate responds too slowly to define a 20- or 30-second effort.
The controller starts with a power estimate, watches the measured response, and makes bounded changes. It records commanded and trainer-reported power separately, but the shipped learning path currently uses trainer-reported overshoot to classify high-range evidence; that is an open defect. When the trainer is already at its floor, the app labels that limitation instead of reporting clean HR-control authority.
What the software evidence says
| Claim | Evidence | Result | Verdict |
|---|---|---|---|
| Bounded safety behavior | Mode-specific checks and out-of-family simulations | No new max-HR breach scenarios appeared in the out-of-family set, but the current Z4 sweep still has supervisor stacking and release failures | Software evidence with open gaps |
| Bounded HR-guided control | Controller checks plus one recorded hardware calibration ride | The hardware ride executed cleanly, but the envelope still contains low-gain late/never arrival and Z4 supervisor failures | Partial hardware evidence |
| Learns rider cardiac gain | Synthetic rides plus the first recorded hardware calibration ride | The synthetic result did not cover the shipped failure: one trainer overshoot contaminated and persisted the hardware rider's high-range model | Not supported in the shipped path |
| Cross-workout knowledge transfer | Interval-to-Zone-2 simulation | A historical simulation produced 58.5% lower feedforward error; it does not validate the model currently persisted by the app | Simulation only |
| First-ride calibration quality | Simulation plus the first recorded hardware calibration ride | The hardware ride controlled cleanly but wrote a corrupted high-range model from one below-gate sample | Not supported |
| Compensates for cardiac drift | Cardiac drift simulation | Drift MAE 0.071 bpm/min and compensated HR error 0.65 bpm in the synthetic model | Software evidence only |
Most evidence here is software, synthetic, computational, or replay-based. The first recorded hardware calibration ride adds one narrow observation: control execution passed on that ride, while learning failed. None of this is clinical validation, a controlled rider study, or a medical-device evaluation.
System Architecture
The system uses a simple first-order approximation of cardiovascular response: heart rate at a given measured power is represented by a personal baseline plus a gain factor (K) multiplied by power, with a response delay. The controller uses that estimate for its starting point, clips it to hardware limits, then applies small feedback corrections.
In plain terms: the system predicts what power should put you in zone, asks the trainer for that resistance when the trainer can execute it, and makes small adjustments based on what your heart rate actually does. Lower model error should improve the first guess, but the current persisted learning path has not earned that product claim.
The Core Equation
Where K is the model's cardiac-gain estimate (bpm per watt), β0 is resting baseline, δ is the drift estimate, and K_safe is a separate bound intended to bias toward lower power when uncertainty is high. On the recorded hardware calibration ride, that safety ceiling remained conservative even though the learned feedforward model was wrong.
Control follows the workout timescale
The controller selects different behaviors depending on workout phase and effort duration. This matters because heart rate lags power changes, and closed-loop HR control only makes sense when the effort lasts long enough for the cardiovascular system to respond.
| Mode | When Active | Strategy |
|---|---|---|
| Open Loop Warmup | First 150s + ramp + dither | Power-controlled. HR observed for estimator but not used for control. |
| Calibrating | Initial K estimation | Ternary dither excitation. Estimator running, feedforward not yet active. |
| Feedforward + Trim | Longer steady work | K-based power estimate + bounded HR trim. |
| Continuous HR | Supported lower-zone steady work | Slow filtered HR feedback with bounded output and a separate hard max-HR stop. |
| Power-Led Intervals | Short and hard efforts | Power targets lead; HR is a guardrail, not an instant target. |
| Open Loop Recovery | Recovery between intervals | Fixed low power. HR drops naturally without feedback chasing. |
| Power Assist Sprint | Short intervals (<30s) | Power-led targets. HR used only as ceiling. |
Bayesian K Estimation
A bank of parallel first-order models, each assuming a different cardiac time constant, runs simultaneously and is reweighted by predictive fit. The posterior K estimate is a weighted mixture across those models. The coefficient of variation (CVK) influences feedforward authority. The current persistence path does not yet handle contradictory evidence reliably: the first hardware calibration became more confident despite estimates that disagreed across the ride.
What We Tested
The evidence is intentionally scoped. It exercises controller behavior, mode transitions, HR-dropout handling, floor-limited authority reporting, and software safety boundaries. The current sweep also contains open controller failures, so it is not a blanket pass. It does not prove clinical outcomes, broad trainer certification, or real-rider performance under controlled study conditions.
Safety Behaviors
The controller uses escalating intervention layers to reduce overshoot risk. Where predictive safety is active, it checks a model-based lookahead and a slope-based extrapolation, then applies the more conservative limit. A hard maximum-HR stop remains separate.
| Layer | Trigger | Action |
|---|---|---|
| Freeze | Predicted threat exceeds zone_max + 2 bpm | Hold current power; block increases |
| Emergency | Predicted threat exceeds max HR − 5 bpm | Reduce power to 70% of current |
| Hard Stop | Smoothed HR exceeds max HR | Set power to floor |
Evidence Summary
The strongest evidence is narrow: the app can issue bounded trainer commands, label cases where the trainer floor dominates, and avoided new max-HR breach scenarios in one out-of-family test set. Other envelope cells remain red, including Z4 supervisor release, low-gain arrival, and Aerobic Threshold Finder step delivery.
| Behavior | What It Checks | Takeaway |
|---|---|---|
| HR ceiling behavior | Freeze/starvation guard, hard-stop latency compensation, and safety supervisor behavior | Open: Z4 supervisor stays latched below band in 11/15 cells |
| Low-gain arrival | Whether riders with low cardiac gain reach the prescribed band on time | Open: late and never-arrival cells remain |
| Dropout handling | How the controller behaves when HR data becomes stale or unavailable | Falls back conservatively |
| Phase transitions | Work/recovery transitions without feedback chasing or unsafe command jumps | No unsafe command jumps found |
| Trainer limits | Floor-limited rides are reported as limited authority, not clean HR control | Limits are labeled |
| Aerobic Threshold Finder steps | Whether the controller delivers the six prescribed power steps on schedule | Open: 93/100 envelope cells fail step delivery |
| Out-of-family physiology | 0 new max-HR breach scenarios vs. the first-order baseline | No new breach scenarios found |
What the Learning Evidence Shows
The controller's first guess can depend on a stored response-gain estimate (K). Synthetic tests produced promising results, but the first recorded hardware calibration ride falsified the shipped learning path: a brief trainer-reported overshoot was treated as intentional high-range demand, one sample retired the fallback, and the wrong value was persisted. Until that path is repaired and migrated on-device, personalization from learned K is not a supported public claim.
The Protocol Matters
The historical synthetic result was that interval protocols produced lower model error than flat Zone 2 rides. Without meaningful power variation, there is very little information about how heart rate responds to power changes. This does not override the hardware calibration failure described above.
In this synthetic model, intervals produced roughly 10x lower K-estimation bias than flat Zone 2. The test did not reproduce the hardware overshoot and persistence failure.
Cross-Workout Transfer
A historical simulation tested whether useful rides could improve a later controller starting point. In that synthetic sequence, carrying K forward reduced Zone 2 feedforward error by 58.5% compared with starting from a neutral prior. This is a research result, not evidence that the current app persists a trustworthy model.
The synthetic result shows why trustworthy learning could matter. It does not establish trustworthy learning in the shipped app.
Calibration Protocol Design
Historical simulations suggest calibration needs both time and meaningful power variation. The first hardware calibration ride showed that the current protocol and persistence path can still write the wrong high-range model, so those simulations do not support a shipped calibration-quality claim.
Cardiac Drift Compensation
Cardiac drift is the gradual rise in heart rate at constant power during longer exercise. In synthetic drift tests with known ground truth, the drift estimator tracked the simulated drift within 0.65 bpm after compensation. This supports the control concept in software; it is not hardware or outcome validation.
Real-Data Characterization
47 real cycling rides from 5 athletes in the GoldenCheetah OpenData corpus were processed. The K distribution (mean 0.323 bpm/W, range 0.200–0.641) overlaps published literature. The negative correlation between K and peak power (r = −0.72) is directionally consistent with exercise physiology: fitter riders tend to show lower cardiac gain.
Real rides do not provide ground-truth K values, so this is distribution characterization, not accuracy proof. It does not validate the current estimator, evidence gates, or persisted rider model.
Fitness Estimates
Zone Pedal can show an optional VO2max training estimate from its stored cardiac model, using the Storer-Davis cycle ergometry formula applied to an extrapolated maximum power. This is a derived, exploratory estimate, not a measured oxygen-uptake result. Because the current calibration-learning path is not supported by hardware evidence, the estimate should not be treated as validated personalization.
Quality Thresholds
The software attaches a confidence label and withholds the estimate when the stored K posterior is too loose. The first hardware calibration showed that stored confidence can tighten on internally contradictory evidence, so these are software thresholds rather than validated reliability labels.
| CV_K Range | Quality | Estimate Produced? |
|---|---|---|
| < 0.10 | High | Yes |
| 0.10 – 0.20 | Moderate | Yes |
| 0.20 – 0.25 | Low | Yes |
| ≥ 0.25 | Insufficient | No |
VO2max estimation was checked against synthetic ground truth, not against laboratory measurements. No head-to-head comparison with laboratory VO2max testing or other consumer device estimates has been performed.
What We Removed
Several physiological features were built or researched and then kept out of the app. The rule is simple: if the signal cannot change the ride experience reliably, it should not become a rider-facing claim.
| Feature | Decision | Why |
|---|---|---|
| HRV readiness | Not shipped | Moderate recovery states were too hard to distinguish from noise reliably enough for rider-facing advice. |
| Fatigue detector | Removed | The detector missed most simulated fatigue while still producing false alerts under consumer-level noise. |
| Bad-legs-day classifier | Research only | The thresholds are engineering choices, not physiologically derived rider advice. |
Fatigue Detection
The fatigue detector tracked K trajectory changes during a ride. It failed both sides of the tradeoff: 15.0% false positive rate, 15.0% true positive rate, and no useful post-onset detections. Under noise, false positives first exceeded 20% at 4 bpm. That is not good enough to earn a place in the app.
The detector was removed rather than shipped in a state that would train riders to ignore the system.
Stress-Testing Assumptions
Most validation in fitness apps tests the system against its own assumptions. If the estimator assumes a first-order cardiac model, and the simulator uses the same first-order model, then strong results prove internal consistency, not real-world robustness. We also tested against a second simulator that deliberately violates those assumptions, then replayed 52 real cycling rides through the system.
Three Evidence Tiers
Every result in this document falls into one of three tiers:
| Tier | What It Proves | Example |
|---|---|---|
| Exact-match simulation | Internal consistency: the system works when reality matches its assumptions | Synthetic rides with known ground truth |
| Robustness simulation | Measured resilience and failure boundaries when assumptions are violated | Out-of-family cardiac model |
| Real-data replay | Behavioral characterization on recorded rides, not a controlled outcome study | Recorded ride files |
Out-of-Family Model Testing
We built a second cardiac simulator with six physiological effects the controller model does not capture: logistic HR saturation near max, time-varying K and tau (warmup acceleration + fatigue decay), BLE latency jitter and dropout, Student-t noise with ectopic spikes, heat/hydration drift, and a soft HR ceiling. Then we ran the main checks against it.
| Behavior | Count | Meaning |
|---|---|---|
| Maintained | 12 | Performance equivalent to ideal-model testing |
| Degraded | 5 | Measurable loss but still functional |
| Failed | 3 | Falls below acceptable threshold |
Scope of this result: In this out-of-family test set, the safety supervisor added zero new ceiling activations and zero new max-HR breach scenarios. That does not cover every real-world condition. The primary boundary was time-varying K/tau: warmup and fatigue dynamics that the controller model intentionally simplifies.
Real-Ride Replay
52 recorded cycling rides (47 from the GoldenCheetah OpenData corpus + 5 internal Tacx Neo 2T rides) were replayed through the system. Total: 77 ride-hours. This is replay characterization, not a controlled rider study.
| Metric | Value |
|---|---|
| Rides replayed | 52 |
| Total ride-hours | 77 |
| K estimate median | 0.31 bpm/W (matches synthetic distribution) |
| Feedforward RMSE median | 26.3 bpm |
| Feedforward-fit boundary labels | 0 maintained / 19 degraded / 33 breaks |
| BLE delay sensitivity (0→10s) | +5% RMSE (negligible) |
Real rides introduce dynamics the controller model does not capture: nonlinear cardiac responses, autonomic nervous system effects, environmental conditions, and sensor noise. Under the report's feedforward-fit thresholds, breaks means model-fit error above 25 bpm; it does not mean a trainer crash or an HR safety event. The replay set exposes where the simplified model loses fit, while the safety supervisor is assessed separately.
DFA Alpha1 Threshold Detection
DFA alpha1 is an exploratory HRV signal intended for use during a stepped power ramp. The Aerobic Threshold Finder targets a 35-minute protocol with 6 power steps and requires a chest strap transmitting RR intervals. The current controller envelope shows that the workout often fails to deliver those prescribed steps, so its end-to-end threshold result is not currently supported. It is not a lab-equivalent threshold or a medical measurement.
The validation tested four dimensions across 16 criteria:
| Experiment | What It Tests | Result |
|---|---|---|
| Alpha1 accuracy | Computation against known signals (white noise, 1/f, physiological) | All 4 checks met |
| Threshold detection | VT1/VT2 crossing on synthetic ramps (gradual, steep, varying gaps) | All 4 checks met |
| Noise robustness | Detection under typical chest strap noise (σ=10ms, 2% ectopic, 1% dropout) | All 4 checks met |
| Protocol end-to-end | Historic simulation with 5 rider archetypes, compared with the current controller envelope | Open: current sweep fails prescribed-step delivery in 93/100 cells |
The computation and breakpoint checks establish that the analysis algorithm behaves as coded under synthetic inputs. They do not prove that the current workout delivers the required ramp, or that its result agrees with ventilatory or lactate thresholds across riders.
Known Limitations and Honest Gaps
No Clinical Trial
Most results in this document come from deterministic synthetic simulations or Monte Carlo parameter recovery on simulated riders. The out-of-family model testing and real-ride replay extend beyond internal-consistency validation, but they are still computational evidence. No clinical trial has been conducted. No real riders have been studied under controlled conditions with the Zone Pedal controller active.
Current Learning Path Is Not Hardware-Supported
The estimator is intended to infer a response-gain value (K) from power variation during a ride. On the first recorded hardware calibration, trainer ripple crossed a range threshold that commanded power never crossed; one sample contaminated the high-range estimate, retired its fallback, and was persisted. The same ride became more confident despite internally contradictory estimates. Repairs and an on-device migration are required before learned personalization is a supported product claim.
Zone 2 Is Structurally Unobservable
Zone 2 rides provide near-zero Fisher information about K because power variation is minimal. The estimator cannot identify K from flat Zone 2 work alone: without power variation, the HR-power relationship is not identifiable. The software includes a low-variance gate, but that gate does not protect against the high-range contamination observed on the hardware calibration ride.
Current Controller Envelope Has Open Failures
The latest software sweep is not uniformly green. Z4 steady work has supervisor stacking and release failures in 11 of 15 cells, low-gain riders can arrive late or never reach the band, and Aerobic Threshold Finder misses prescribed power-step delivery in 93 of 100 cells. One hardware calibration ride supports clean transport and control execution for that rider and workout only; it does not close those envelope failures.
Fatigue Detection Was Removed
We built a fatigue detector, tested it, and removed it from the app. The single-signal approach could not separate fatigue from sensor noise well enough to trust. We chose to remove the feature rather than ship one that would erode trust in the rest of the system.
Limited Real-World Data
The replay set includes 52 rides from 6 athletes: 47 GoldenCheetah rides plus 5 developer rides. That is enough to characterize replay behavior, but not enough for population-level inference. Elite, elderly, cardiac-compromised, and pediatric populations are not represented.
No Head-to-Head Outcome Comparison
No study has compared training outcomes between Zone Pedal's HR-adaptive control and conventional power-based ERG training. The evidence for HR-based training equivalence comes from Akubat et al. (2013), not from Zone Pedal's specific implementation.
No Medical Device Claim
Zone Pedal is a fitness application. It has not been evaluated by the FDA or any regulatory body. It does not diagnose, treat, or prevent any disease. Users with cardiovascular conditions should consult their physician before using any exercise equipment.
Unvalidated Capabilities
| Capability | Status |
|---|---|
| Learned personalized starting power | Not currently supported by hardware evidence |
| Aerobic Threshold Finder step delivery across the envelope | Open controller defect |
| Broad HR-guided steady-zone delivery | Open envelope defects |
| VO2max estimation against laboratory reference | Not yet validated |
| Long-term fitness trend detection from K trajectory | Not yet validated |
References
Published Literature
Akubat I, Patel E, Barrett S, Sherwin Z. Methods of monitoring training load and their relationships to changes in fitness and performance in competitive road cyclists. J Sports Med Phys Fitness. 2013. PMC3737823.
Argha A, Su SW, Celler BG. Automated PID control of heart rate during treadmill exercise. J Biomech Eng. 2016.
Argha A, Su SW, Celler BG. Heart rate regulation during cycle-ergometer exercise via bio-feedback. Conf Proc IEEE Eng Med Biol Soc. 2017.
Hunt KJ, Fankhauser SE. Heart rate control during treadmill exercise using input-sensitivity shaping. J Sports Sci Med. 2019;18(1):47-55. PMC6370964.
Hunt KJ, Hurlimann N, Fankhauser SE. Physiological systems modelling for heart rate control during cycle-ergometer exercise. Proc Inst Mech Eng H. 2024.
Coyle EF, Gonzalez-Alonso J. Cardiovascular drift during prolonged exercise. Sports Med. 2001.
Achten J, Jeukendrup AE. Heart rate monitoring: applications and limitations. Sports Med. 2003.
Storer TW, Davis JA, Caiozzo VJ. Accurate prediction of VO2max in cycle ergometry. Med Sci Sports Exerc. 1990;22(5):704-712.
Evidence Notes
| Claim | Evidence Used |
|---|---|
| Safety checks | Software checks cover supervisor behavior, dropout handling, phase transitions, floor-limited authority reporting, and out-of-family max-HR behavior; the current Z4 supervisor sweep still contains open failures. |
| Real-ride replay | 52 recorded rides, 76.9 ride-hours, and 26.3 bpm median feedforward RMSE characterize behavior on real files. |
| Cross-workout transfer | A historical simulation showed a 58.5% Zone 2 feedforward-error reduction when synthetic interval-learned K carried forward; this is not shipped-path proof. |
| Calibration learning | The first recorded hardware calibration controlled cleanly but persisted a corrupted high-range model; current learning and personalization claims are not supported. |
| Threshold workout | Algorithm checks pass on synthetic signals, but the current controller sweep fails prescribed Aerobic Threshold Finder step delivery in 93/100 cells. |
| Fatigue detector decision | 15.0% false-positive rate, 15.0% true-positive rate, and no useful post-onset detections led to removal. |