Skip to main content
Running Power Across Garmin, Apple Watch, and Stryd: A Careful Field-Checking Guide
Training & Performance ·

Running Power Across Garmin, Apple Watch, and Stryd: A Careful Field-Checking Guide

Compare running-power trends across Garmin, Apple Watch, and Stryd without treating any proprietary estimate—or one field test—as laboratory truth.

SensAI Team

16 min read

SensAI

Get a training plan that adapts to your recovery

Download on the App Store

Running Power Across Garmin, Apple Watch, and Stryd: A Careful Field-Checking Guide

If you want the short answer: treat running power as a proprietary estimate, not laboratory truth, and check whether one device is repeatable enough to support your training decisions.

The sessions below are coaching checks, not a validated cross-device protocol. They can help you notice obvious inconsistency between power, pace, heart rate, perceived effort, terrain, and conditions. They cannot prove that a watch or footpod measures true mechanical or metabolic power.

Why running power numbers differ across Garmin, Apple Watch, and Stryd (and why none is universally “right”)

Different devices can be useful and still disagree. That is normal.

Ray Maker (DC Rainmaker) put it directly: “There is no agreed-upon scientific standard for running power.”1

Garmin, Apple, and Stryd each use different sensor stacks, assumptions, and smoothing choices. Garmin also notes that your default power zones may not match your personal ability until you customize them.2 So the right question is not “Which device is objectively correct?” It is “Is one device repeatable and interpretable enough to support consistent training decisions?”

A practical approach is to judge repeatability and usefulness for a defined purpose, not to expect equal watts across manufacturers.

Model inputs, elastic recoil assumptions, wind/grade handling, and sensor source differences

At a high level, differences come from four buckets:

  1. Model inputs: cadence, speed source, barometer/GPS quality, wrist motion, accessory sensors.
  2. Biomechanical assumptions: how each model treats elastic recoil and center-of-mass mechanics.
  3. Environment handling: wind and grade compensation vary by platform and configuration.3
  4. Signal source placement: wrist-only vs footpod vs chest + watch combinations change noise and lag profiles.

That last point matters. In a 2021 comparison of five running-power technologies, Stryd showed the strongest repeatability (SEM <=12.5 W, CV <=4.3%, ICC >=0.980) and best concurrent validity to VO2 among tested devices in trained runners.4 As Víctor Cerezuela-Espejo and colleagues wrote, “Stryd device was found as the most repeatable technology… besides the best concurrent validity to the VO2.”4 But “most repeatable in this study” is still not “always right in every athlete, every terrain.”

Pre-test setup checklist to reduce noise before field checking

Before you test, reduce avoidable error. Inconsistent setup can make an already model-dependent estimate harder to interpret.

Firmware, body-mass accuracy, sensor pairing, wind settings, route selection, footwear consistency

Use this checklist before Session A:

  • Update watch/footpod firmware (Garmin, watchOS, Stryd app).
  • Confirm body mass is current (power estimates are mass-sensitive).
  • Use one speed source across a series of comparable checks (do not mix treadmill GPS drift with open-sky GPS).
  • Verify accessory pairing and battery state.
  • Confirm wind setting behavior on the platform you use (especially Garmin running power options).3
  • Choose repeatable routes: flat loop for steady test, consistent hill grade for repeats.
  • Keep footwear consistent across all sessions.
  • Control timing, heat exposure, and caffeine so HR drift interpretation is cleaner.5

If setup quality is poor, postpone interpretation rather than forcing precision from noisy data.

Step 1 — Identify the device-specific setting you are reviewing

Before changing zones, identify how the selected ecosystem defines its threshold or critical-power anchor. A value from one system is not automatically transferable to another.

Garmin TP/FTP setup and zone reset logic

On compatible Garmin devices, running-power zones can be based on a manually entered threshold-power value or on watts, and the zones can be reset to their defaults.2 Follow the instructions for your specific watch because menus and supported features differ.

Practical check:

  • Confirm which threshold-power value the watch currently uses.
  • Change that value only when you have a defensible device-specific estimate or qualified coaching guidance.
  • After changing it, confirm that the displayed zones match the intended percentages or watt boundaries.

Apple Watch baseline capture and workout-view configuration

Apple Watch supports running power and customizable workout metrics; make sure power is visible in your workout view before testing.6

For baseline capture:

  • Use the same watch, band tightness, and wrist placement each time.
  • Record power, pace, HR, cadence, and RPE together.
  • Avoid making threshold calls from one run; if you perform the optional field checks below, consider them together.

Apple Watch displays running power, but Apple does not present that metric as the same quantity as Stryd Critical Power or Garmin threshold power. Use repeated sessions to understand its behavior rather than importing a threshold from another device without checking.

Stryd CP pathways (17-min estimate, manual test, and 90-day model hygiene)

Stryd gives three practical paths:7

  1. Model-estimated CP from recent training.
  2. 17-minute estimate with two 60-second surges.
  3. Manual maximal test inputs.

Stryd also states CP uses roughly a 90-day window and recommends max-effort testing at least every 90 days.7 That means your CP can change even when your “fitness feeling” is flat.

Useful context from the literature:

  • At submaximal speeds, Stryd power correlated strongly with oxygen consumption (R²=0.82) and external mechanical power (R²=0.88) in recreational runners.8
  • Spatiotemporal reliability is strong for key metrics (CV <3% for most variables), though some metrics like flight time are less stable.9
  • Felipe Garcia-Pinillos and colleagues concluded the pod is practical and that it provides accurate step length/frequency while underestimating contact time and overestimating flight time, which is relevant when you interpret form metrics beside power.9

Step 2 — Optional field checks (steady run + hills + threshold effort)

If these sessions are appropriate for your training and health, place them far enough apart to recover and keep conditions as comparable as practical. This three-session sequence is an illustrative checking framework, not a scientifically validated minimum or timeline.

Session A: Steady-state power/pace/HR drift audit

Goal: observe whether power, pace, heart rate, and perceived effort behave consistently during one steady session.

Protocol:

  • 15-minute easy warm-up.
  • 40-50 minutes steady in upper easy/lower moderate domain.
  • Hold a conservative effort or device-power target reasonably steady; observe pace, HR, and perceived effort.

Interpretation:

  • If HR and RPE climb at fixed device power, heat, hydration, fatigue, pacing, terrain, and sensor behavior are all possible contributors. The session does not identify which one.5
  • In a controlled heat study, cardiovascular drift during running included a large HR rise and stroke-volume decline. That laboratory result illustrates why environmental context matters; it is not a universal correction factor for field thresholds.5
  • Jonathan E. Wingo and colleagues summarized the mechanism clearly: “The upward drift in heart rate associated with CV drift reflects increased relative metabolic intensity.”5

Session B: Hill repeat consistency and downhill decoupling check

Goal: observe terrain sensitivity and short-repeat consistency.

Protocol:

  • Use a familiar number and duration of uphill repetitions that already fit your program; do not add maximal hill work solely to test a device.
  • Jog down easy.
  • Track rep-to-rep power dispersion and RPE stability.

Interpretation:

  • If power is highly unstable while pace and effort are stable, suspect device/model terrain bias.
  • Check downhill behavior: some models decouple oddly during eccentric-biased running.

There is no validated dispersion threshold in this guide. Treat large, repeatable terrain-specific differences as a reason to inspect the device setup and manufacturer documentation, not as proof of physiological error.

Session C: Device-specific threshold effort, when appropriate

Goal: obtain a device-specific estimate using a protocol you understand and can perform safely.

One coaching option is a 3-minute plus 9-minute maximal-effort test with full recovery, followed by the calculation described by TrainingPeaks.10 This is a secondary coaching protocol, not a peer-reviewed validation of Garmin, Apple, or Stryd watts. Maximal testing is not appropriate for everyone; use medical clearance or qualified supervision when your health or training history makes that relevant.

Another research protocol is the 3-minute all-out test (3MT). In eight physically active participants running tethered on a non-motorized treadmill, end-test power was similar to modeled critical power, while the anaerobic work estimate was not valid.11 That narrow laboratory finding does not validate a wrist or footpod estimate in ordinary outdoor running.

A separate study of eight male endurance runners found critical speed close to intermittent maximal lactate steady-state speed, while continuous MLSS was lower.12 That result concerns running speed and a specific intermittent protocol—not commercial-device power or an interchangeable run FTP.

Step 3 — Cross-signal review when power conflicts with pace, HR, and RPE

Power is one signal. Good coaching needs signal arbitration.

Keep: the current setting remains useful

Keeping the current device-specific setting may be reasonable when:

  • Session A drift behavior is stable.
  • Session B rep consistency is acceptable.
  • A device-appropriate threshold estimate is broadly consistent with prior comparable testing.
  • Pace, HR, and RPE broadly agree with the power story.

These observations support continuity; they do not establish a formal confidence score or measurement validity.

Review: persistent mismatch with a plausible device or setup cause

Review the setup when the mismatch is systematic rather than random. Examples:

  • One device reads consistently high on hills vs perceived effort.
  • Threshold sessions repeatedly feel one full zone harder than prescribed.
  • Power-pace relationship is stable but shifted (a possible device or model effect).

If you decide to adjust a device-specific threshold, change one variable at a time:

  1. Update threshold anchor (not every zone manually).
  2. Regenerate zones.
  3. Re-check with one confirmation session.

Repeat later: conditions or execution compromised the check

Repeat only when another hard test fits the program and the original check was materially compromised:

  • Unusual heat/humidity, poor sleep, or heavy residual fatigue.5
  • Inconsistent route conditions.
  • Pacing errors in Session C (especially early overpacing).

There is no formal low-confidence label or automatic repeat window in SensAI. The runner or coach decides whether another test is warranted.

Using device-specific zones after field checks (different display, shared training intent)

Your devices can display different watts while still supporting the same training intent.

Set zones per device, but keep the intent map shared:

  • Endurance: low metabolic strain, conversational breathing.
  • Tempo: durable moderate strain, controlled breathing.
  • Threshold: hard but repeatable, limited-talk effort.
  • VO2/anaerobic: short severe work, strict recoveries.

This reduces “zone identity crisis” when switching between Garmin, Apple Watch, and Stryd.

Converting a device-specific estimate into endurance, tempo, or threshold intervals

If you use a device-specific threshold estimate, keep interval prescriptions inside that ecosystem and review how the sessions actually feel and perform:

  • Endurance sessions: conservative fraction of threshold, long duration.
  • Tempo sessions: moderate fraction, extended repeats.
  • Threshold sessions: near-threshold repeats with controlled recoveries.

The percentages and terminology vary by ecosystem. Use the model that belongs to the selected device or coaching system, and review response over time rather than calling the threshold “validated.”

Why running power zones can change (fitness, model window, environmental load)

A zone change can reflect one or more of four things:

  1. Fitness changed (up or down).
  2. Model window rolled (e.g., Stryd’s 90-day behavior updated CP).7
  3. Environment shifted (heat/wind/terrain seasonality).
  4. Data quality changed (new shoes, firmware, sensor placement).

A 2025 systematic review of field-based critical-speed tests screened 450 records and included 19 studies. It found that time-trial and 3MT approaches can be reliable under specified conditions, while protocol details matter.13 It did not validate commercial running-power meters or prove that device zones should change every week.

What SensAI can—and cannot—do with running data

SensAI can use compatible aggregated HealthKit workout and recovery data as personal context for its LLM coach. Apple Watch connects directly; compatible Garmin metrics can flow through HealthKit. Stryd support is not part of the current product contract. SensAI does not automatically ingest or normalize Garmin, Apple, and Stryd running-power estimates as a cross-device calibration layer.

  • The coach can discuss the power, pace, heart-rate, perceived-effort, and terrain context you provide.
  • Aggregated recovery and completed-workout context can inform weekly program regeneration.
  • The app does not assign High/Moderate/Low validation tiers, certify a device, or automatically rewrite its power zones.
  • During a workout, changes occur when you request them through conversation or a quick action.

The useful coaching question remains: is one consistent setup helping you pace and review training without creating false precision? Keep device-specific watts within that device’s ecosystem unless you have evidence supporting a conversion.

FAQ quick answers mapped to target queries

Is Garmin running power accurate?

Garmin running power can be repeatable and useful, but there is no universal ground-truth standard for consumer running power. Use the correct setup for your model, keep the source consistent, and assess repeatability for your intended use.24

How do I find critical power for running?

Use the method defined by the ecosystem you plan to train with. Stryd offers model-based, 17-minute estimate, and manual pathways.7 Other coaching systems use different critical-speed, threshold-power, or run-FTP methods; do not assume the outputs are interchangeable.

How do I set running power zones on Garmin?

On a compatible watch, open running power zones, select threshold power or watts as the basis, enter a threshold value if known, and adjust or reset the zones as needed.2 Follow the manual for the exact model.

Apple Watch running power zones vs Stryd: which should I trust?

Use one system consistently for a defined purpose. Cross-device divergence is expected because the manufacturers use different proprietary estimates and there is no agreed conversion standard.1

Running power vs heart rate for threshold training: which wins?

Neither alone. Use power for workload targeting and HR for strain context. In heat, HR drift can rise materially over time at fixed workloads, so single-signal decisions can misclassify intensity.5

What is a good critical power test protocol for runners?

There is no single best protocol for every runner or device. Follow a device-specific method or qualified coaching protocol. Critical-speed and tethered-running studies can inform physiology, but they do not validate consumer-device watts.121113

Why do running power zones keep changing?

Possible reasons include a rolling model window, new maximal efforts, a changed manual threshold, device settings, firmware, body-mass entry, or data quality. In Stryd, the model uses roughly 90 days of running data.7 A zone change does not prove that fitness changed.

How do I validate running power with pace and heart rate?

Use a cautious review sequence:

  1. Run standardized sessions.
  2. Compare power to pace/HR/RPE behavior.
  3. Keep the setup when it remains useful, inspect persistent mismatch, and repeat hard testing only when appropriate.

This is a coaching heuristic, not a validated SensAI workflow.

Continue with SensAI

If you remember one line, use this one: running power can be useful without being laboratory truth. Keep the device consistent, understand its model, and avoid converting one proprietary watt value directly into another.


References

Footnotes

  1. Maker R. “Apple Watch Running Power Data Comparison vs Garmin/Stryd/Polar/COROS.” DC Rainmaker, 2022. https://www.dcrainmaker.com/2022/06/running-comparison-garmin.html 2

  2. Garmin. “Forerunner 265 Owner’s Manual — Setting Your Power Zones.” Garmin Support, accessed 2026-07-12. https://www8.garmin.com/manuals/webhelp/GUID-F41EAFB3-6CC9-42DE-9C6C-9E358DBB0671/EN-US/GUID-28DE6904-5F2F-47B9-AD8C-BCF3F5FE445E.html 2 3 4

  3. Garmin. “Forerunner 255 Owner’s Manual — Running Power.” Garmin Support, accessed 2026-07-12. https://www8.garmin.com/manuals/webhelp/GUID-676967A0-1B23-4384-9BC9-76F3D643F1C8/EN-US/GUID-D74FC870-3A94-4376-81D5-C9484545EAD9.html 2

  4. Cerezuela-Espejo V, et al. “Are we ready to measure running power? Repeatability and concurrent validity of five commercial technologies.” European Journal of Sport Science, 2021. https://pubmed.ncbi.nlm.nih.gov/32212955/ 2 3

  5. Wingo JE, et al. “Cardiovascular Drift and Maximal Oxygen Uptake during Running and Cycling in the Heat.” Medicine & Science in Sports & Exercise, 2020. https://pubmed.ncbi.nlm.nih.gov/32102057/ 2 3 4 5 6

  6. Apple. “Workout views and running metrics on Apple Watch.” Apple Support, accessed 2026-07-12. https://support.apple.com/guide/watch/workout-views-and-running-metrics-apd1f24d4d35/watchos

  7. Stryd. “Critical Power Definition.” Stryd Help Center, 2025. https://help.stryd.com/en/articles/6879345-critical-power-definition 2 3 4 5

  8. Imbach F, et al. “Validity of the Stryd Power Meter in Measuring Running Parameters at Submaximal Speeds.” Sports, 2020. https://pmc.ncbi.nlm.nih.gov/articles/PMC7404478/

  9. García-Pinillos F, et al. “Absolute Reliability and Concurrent Validity of the Stryd System for the Assessment of Running Stride Kinematics at Different Velocities.” Journal of Strength and Conditioning Research, 2021. https://pubmed.ncbi.nlm.nih.gov/29781934/ 2

  10. TrainingPeaks. “Running With Power: How to Find Your Run FTP.” TrainingPeaks Learn, accessed 2026-07-12. https://www.trainingpeaks.com/learn/articles/running-with-power-how-to-find-your-run-ftp/

  11. Gama MCT, et al. “The 3-min all-out test is valid for determining critical power but not anaerobic work capacity in tethered running.” PLOS ONE, 2018. https://pmc.ncbi.nlm.nih.gov/articles/PMC5812641/ 2

  12. de Lucas RD, et al. “Is the critical running speed related to the intermittent maximal lactate steady state?” Journal of Sports Science and Medicine, 2012. https://pmc.ncbi.nlm.nih.gov/articles/PMC3737850/ 2

  13. Lipková L, et al. “Field-based tests for determining critical speed among runners and its practical application: a systematic review.” Frontiers in Sports and Active Living, 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC11933073/ 2

SensAI

SensAI

Free AI fitness coach

Get Free