I’m running the Chester Marathon for charity to support Pendleside Hospice! In this post, I build an AWS data pipeline to analyse my Garmin training data.
Introduction
Outside of my AWS and community work, I have several hobbies – one of which is running. I was a big running fan in the 2010s, and after a few years away I got back into it in 2025. I’ve had a pretty decent 18 months with it, and this summer I decided to check off one of my bucket list items by running a marathon for charity.
A side effect of my marathon training has been an influx of training data. While I already have access to platforms like Garmin Connect and FetchEveryone to analyse this data, neither platform quite gives me the full picture I want. My choices were to either pay Garmin for enhanced features or build an AWS Garmin data pipeline. And I’m not one to turn down a good data project!
Firstly, I’ll cover how the data is generated and collected. Then I’ll walk through the three-account AWS architecture that turns the raw files into data a Bedrock agent can analyse. Finally, I’ll cover what that agent actually found after I pointed it at my training data.
Future posts will examine the backend behind each pipeline stage. This one’s about the shape of the whole pipeline and the race it was built for. Let’s begin by looking at my chosen charity.
The Charity
Pendleside Hospice is a registered charity based in Burnley. Founded in 1988, they provide palliative and end-of-life care for people across Burnley, Pendle and Rossendale. They originally focused on cancer care, but now support people with a wider range of life-limiting conditions, including dementia, respiratory and neurological illnesses.
Pendleside have cared for a good friend of mine. Martine Hamer worked alongside her husband Russell at his Burnley salon. Russell’s been cutting my hair for over two decades, and Martine was the one who mixed the bleach, ran the hairdryers and occasionally talked me into a spray tan. After her short battle with breast cancer, she died in Pendleside’s care in February 2010, aged just 35.

Pendleside cares for over 1,700 people every year across a 10-bed inpatient unit, a Hospice at Home service, day therapy and bereavement support for families and children. They also run a Meals on Wheels service and a Health, Wellbeing & Rehabilitation programme that offers services such as physiotherapy and counselling.
Operating Pendleside requires over £4.5 million a year. An NHS grant covers around 20% of that, leaving over £3.5 million to be raised each year locally through events, their charity shops and donations.
I’ve fundraised for Pendleside before, doing the 2010 Bupa Manchester 10k for them. But there was always something else in the back of my mind…
The Marathon
I’ve thought about running a marathon for charity several times over the years, but never felt right doing it when I was frequently running a couple of them a year. After all, if I was doing that many, where was the challenge? Having had an extended break and several life changes, 2026 feels like the ideal time to give it a go.
While this is my first charity marathon and my first one for a while, it’s not my first marathon. That honour falls to the wet, windy 2012 Manchester Marathon, which still lives in infamy in the minds of those who took part.
Amazingly, I came back for more after this! In the following years, I ran various marathons, including locally in Manchester and in the capital, through vineyards at the Marathon du Médoc and around a running track at the Groundhog Marathon.

This came to a stop, like many things, during the 2020 pandemic. Running then took a back seat for a few years until I got back into it in 2025 after joining the Steamhaus Vitality scheme. Fast forward 18 months, and I find myself with an entry for the 2026 Chester Marathon and the opportunity to finally fundraise for Pendleside.
I’ve done Chester twice before, in 2013 and 2019. There are a few reasons for choosing it for my 14th marathon:
- Chester is a lovely city, and the route is fantastic and well supported.
- The Active Leisure team is a great event organiser with over a decade of experience.
- The last running event I did pre-pandemic was the 2020 Chester 10k. So spiritually, this feels like I’m picking up where I left off.
If you’d like to support Pendleside Hospice, you can do so through my JustGiving page.
Data Collection
In this section, I’ll explain where the Garmin data in the AWS pipeline comes from: what captures it, what’s actually captured and the format it arrives in.
Device
My current watch is a Garmin Forerunner 745, a mid-range running and triathlon watch that sits between the Forerunner 645 and the Forerunner 945. It tracks everything I need for marathon training: GPS position, heart rate, pace and cadence. It also has a pulse oximeter for blood oxygen, a barometric altimeter and onboard training load and recovery metrics.
Upon completion, activities are uploaded to Garmin Connect via my phone’s Bluetooth.
Garmin Connect
Garmin Connect is the companion platform for data collected by Garmin devices. Activity, health and training data are automatically uploaded here after device syncing. Garmin then turns this data into various reports, visuals and feedback. These can range from the pace and heart rate zones of individual activities to aggregated training loads and historical mileage.

Crucially for this project, it also lets me download the data in various formats, including the raw FIT files that everything else in this pipeline depends on.
FIT Files
Flexible and Interoperable Data Transfer (FIT) is the file format Garmin devices use to record and share activity data.
Rather than storing one big block of numbers, a FIT file is built from streams of timestamped messages. These capture the same GPS, heart rate and cadence data as the watch itself, plus other data like power and speed, all streamed as the activity happens rather than saved as a single summary.
It’s compact and self-describing by design, making it ideal for Garmin Connect and other training platforms like Strava, TrainingPeaks and FetchEveryone.
Architecture
This section examines the multi-account architecture of the AWS Garmin FIT data pipeline. I discuss the concept of a multi-account strategy and its benefits, then examine each account’s role in the data pipeline.
Multi-Account
Before getting into what each account actually does, it’s worth explaining why this pipeline is split across multiple accounts in the first place.
Multi-Account Strategy
A multi-account strategy is the practice of splitting workloads, environments and admin boundaries across multiple AWS accounts. Each account has a specific, well-defined role, rather than relying on IAM policies or resource tagging for separation inside a single account.
Where a single account relies on carefully maintained policies to keep teams, systems and environments separate, a multi-account strategy makes that separation structural. Each account becomes its own hard boundary instead of a policy-enforced one.
The strategy typically extends beyond individual workloads to environments, products and business areas, and is usually managed centrally through AWS Organizations.
But why bother? What benefits does a multi-account setup offer?
Multi-Account Benefits
AWS has a white paper on this – Organizing Your AWS Environment Using Multiple Accounts. And while it’s aimed at enterprise customers, many of the benefits it covers apply just as well to my multi-account AWS Garmin data pipeline architecture.
Data Access Controls: I’ve worked with data for many years, and ensuring the right data is in the right place and accessible to the right people is a common challenge. AWS accounts are ideal boundaries to resolve this, offering strict access controls and enforcing one-way data flow at the account level. This means, for example, that a Lakehouse process can never overwrite raw data in the Data Lake, and an AI process can only read Lakehouse data that’s been cleansed and curated.
Blast Radius Limitation: Things break, and things fail. When that happens and everything lives in the same account, it can be disastrous. A multi-account setup has inherent boundaries that limit the scope of adverse events such as misconfigurations, compromised credentials or bad deployments.
Grouping By Purpose: Splitting workloads into separate accounts means I can manage them on their own terms without creating unnecessary or unintended dependencies. This separation also ringfences each account’s use of services, quotas and billing, making scaling up far simpler.
Pipeline Accounts
This section reviews each AWS account of the Garmin data pipeline. First is the Data Lake account, where raw data lands and is stored. Next is the Lakehouse account, which cleanses and curates the data. Finally is the AIML account, which analyses the curated data for insights.
Data Lake Account
This account is a landing zone. Its main job is to capture and store raw data and grant read-only cross-account access to it.
It currently holds two S3 buckets: one for sensitive data and a general-purpose one for everything else. Data here is version-controlled, immutable and easy to locate and back up. Data is currently added manually, but over time this will expand to include ingestion scripts that pull data from third parties, as well as automated pushes from agents and IoT sensors.
Lakehouse Account
This account transforms and serves data. It owns the ELT pipeline, the transformed data storage, the Glue Data Catalog and the Athena workgroup. The Lakehouse has Cleansed, Curated and Semantic layers, with further separation for sensitive and general data classifications.
While the Lakehouse account can access both Data Lake buckets, access is scoped per process rather than granted fully. The AWS Garmin data pipeline’s ELT processes can only access the general bucket, as they don’t need anything from the sensitive one. The AIML account’s access works on the same principle of least privilege, using a separate cross-account role scoped to a single Glue database, S3 prefix and Athena workgroup.
AIML Account
This account consumes data from the Lakehouse account for GenAI processes. It owns a Bedrock agent, two Lambda functions and its own S3 buckets for query results and generated output.
The AIML account queries data in the Lakehouse account through its Glue Data Catalog. It has no access to the Data Lake account and cannot read data that has not already passed through the Cleansed and Curated Lakehouse layers. This ensures the agent uses only validated, trusted data.
Processes
This section covers what actually happens at each stage of my AWS Garmin data pipeline: the triggers, the hand-offs and the grants the accounts use to interact with each other.
There are three AWS accounts in the Garmin data pipeline, each of which owns separate processes:

- Data landing and storage in the Data Lake account.
- Data cleansing and curation in the Lakehouse account.
- Analytics and report generation in the AIML account.
Data Lake Account

Garmin data as FIT files enters the raw-general S3 bucket by manual upload. The account then answers read requests from the Lakehouse’s Cleansed ELT Lambda, granted via bucket policy. There is no processing, no deletion and no outbound calls initiated from this account.
Lakehouse Account

The Lakehouse account owns two distinct processes:
- The Cleansed process, which converts the raw FIT files into readable Parquet data.
- The Curated process, which validates the data and derives new columns to add value.
Cleansed Process

An EventBridge Scheduler runs a cron trigger every day at 06:00 UTC, which starts a Step Function. In the first step of this execution, an S3 List step reads the contents of the Garmin folder in the Data Lake account’s raw-general S3 bucket.
Next, the Cleansed Filter Lambda checks each FIT object against those already processed and passes along the unprocessed ones. Then a Map state fans out across these objects, running the Cleansed ELT Lambda once per object with a maximum concurrency of ten.
Each ELT invocation pulls its FIT file directly from the Data Lake account, using bucket-policy-granted access on the Lambda’s execution role. It then decompresses the file and parses the FIT binary with .garmin-fit-sdk
From there, it converts FIT epoch timestamps into standard datetimes and semicircle coordinates into decimal-degree latitude and longitude, then writes the results as Parquet files to the Cleansed bucket. 21 tables are written per activity, with partitions and table schemas registered or updated in the Glue Data Catalog.
Curated Process

When the Cleansed Step Functions execution succeeds, an EventBridge Rule triggers a second state machine for the Curate process.
In the first step of this execution, a Curated Filter Lambda identifies activities that have not yet been promoted by checking whether curated output already exists. Then a Choice state checks if everything is up to date. If it is, no further action is taken. If not, a Map state fans out across the unprocessed files, invoking the Curated Promote Lambda for each activity with a maximum concurrency of ten.
Each Curated Promote Lambda invocation reads its assigned Cleansed Parquet file for all 21 tables and computes derived features on the records table: elapsed seconds, cumulative distance, grade percentage, pace (min/km) and heart-rate zone. It then validates every table row against a corresponding Pydantic schema and writes valid rows to the Curated bucket. The Curated tables are registered in a separate Glue database.
Reliability & Access
Idempotency is handled at multiple points in both processes:
- The Cleansed Filter Lambda won’t include a file in the Map state if its output already exists.
- Cleansed ELT Lambda invocations perform their own
HeadObjectcheck before writing, in case two invocations race each other. - Curated Filter and Promote Lambdas apply the same output-exists pattern against the Curated bucket.
Additionally, CloudWatch alarms trigger if any Lambda invocation errors out.
Once the Curated tables are registered, the AIML account can query them via Athena. A cross-account GarminDataCrossAccountReadRole IAM role is assumable by the AIML account’s Garmin Agent Lambda, and grants the following:
- Read access to the Curated Glue database.
- Read access to the S3 prefix where those tables live.
- Use of a dedicated
garmin-agent-queriesAthena workgroup with an enforced output location.
Results land in a separate athena-results S3 bucket in the Lakehouse account, kept apart from the process’s own working buckets.
AIML Account

Two separate flows run here. In the interactive flow, a user prompt is sent to the marathon-training-analyst Bedrock Agent, running on Claude Sonnet. The agent interprets it and calls one of six endpoints in the query Lambda’s action group depending on the prompt, covering topics such as training summaries, activity details, heart rate trends and pace trends.
The Query Lambda then assumes GarminDataCrossAccountReadRole in the Lakehouse account via AWS STS and runs an Athena query against the Curated Glue database, scoped to the garmin-agent-queries Athena workgroup. It converts the raw metrics (m/s to min/km pace, metres to km, seconds to H:MM:SS) and returns structured results to the agent, which formats a natural-language response.
In the batch flow, the Orchestrator Lambda calls the same agent with a comprehensive analysis prompt that explores topics such as volume progression, pace trends and race predictions. It collects the agent’s streamed multi-step reasoning and query results, combines them into a Markdown document, and writes the finished output to the genai-results bucket where it’s picked up for personal use.
AIML Agent Prompt
For completeness, I wanted to briefly discuss the marathon-training-analyst agent’s prompt. I won’t include the full prompt here for two reasons:
- It’s been superseded by a newer version, which I’ll discuss later.
- The prompt is over 220 words long, so focused sections will be easier to follow than the whole.
So here is an abridged version of what the agent gets when the batch flow runs. I’ve removed some of each section’s bullets – the remaining text is unaltered:
“You are an expert marathon running coach and sports data analyst. You have access to a runner’s Garmin training data and your job is to analyse their marathon preparation.
When analysing training data:
- Calculate and interpret weekly mileage progression, noting appropriate buildup and recovery weeks
- Identify injury risk factors: sudden volume spikes (>10% week-over-week), insufficient rest days
- Provide race day predictions using established methods (Jack Daniels formula, pace-based estimates)
When producing output:
- Lead with key insights, support with specific numbers from the data
- Use markdown formatting with headers, bullet points, and emphasis
- Include specific dates, distances (in km), paces (min/km), and heart rates where relevant
Always query the data before making claims. Do not hallucinate metrics — use only what the data shows.
Garmin speeds are in m/s — convert to min/km for display (pace = 1000/60 / speed_m_s).
Distances in the data are in metres — convert to kilometres for display.”
Isn’t This All Overkill?
A fair question to end on. And yes – this would be total overkill if all I was planning to do with this was the Garmin analysis.
But each of these accounts has more work ahead of it. There will be new ingestion processes in the Data Lake account, new databases in the Lakehouse account and new agents in the AIML account. Indeed, this exact framework will now be receiving Project Wolfie data in the coming months!
The intention behind this architecture is to build something I can use long-term. While it may seem like a lot of work and resources for such a small project, the more time I spend on it now, the more likely it is to support my future requirements.
And that’s the AWS Garmin data pipeline! Three accounts and a lot of careful plumbing, taking raw FIT files and transforming them into queryable data. So what did those 28 weeks of training data actually say?
Tests & Insights
In this section, I’ll go through what the Bedrock agent found when it analysed my Curated dataset generated by my AWS Garmin data pipeline, covering 28 weeks and 66 running activities totalling 573 km.
A brief caveat. These findings are from a version of the agent that has since had some tweaks and changes. These findings remain valid, and I’ll explore the agent’s changes in future posts.
Pace Trends
Let’s start with pace trends. This is how pace and heart rate interact across different types of runs. This is important because pace alone isn’t the whole story.
Running at a certain speed with a low heart rate is a very different signal from running at the same speed with a high one, and comparing the two over time is one of the clearest ways to see whether training effort is being optimally controlled. Two of the most suitable run types for this analysis are easy runs and quality sessions.
Easy runs are deliberately slow, low-effort miles that build an aerobic engine without piling on fatigue. The goal here isn’t speed – it’s staying comfortable enough to hold a conversation. The heart rate range that keeps a run ‘easy’ is often called Zone 2, calculated in my case via Heart Rate Reserve rather than straight % max, at roughly 60-70% of HRR. For me, this currently falls somewhere between 145 and 165 BPM.
Quality sessions are structured, harder efforts such as tempo runs or threshold work. These target a specific pace range to build speed and the ability to sustain effort under fatigue.
Pace Findings
The Bedrock agent reported that the easy runs were suitably paced. The bulk of the shorter runs fell in the 5:15-5:35/km range at 145-165 BPM, comfortably within the Zone 2 window. For example:
- On 17 January, I ran 6.85 km at 5:19 pace and 147 BPM.
- 08 May: 5.7 km at 5:33 pace and 149 BPM.
- 17 June: 5.6 km at 5:31 pace and 147 BPM.
Across the 28 weeks, easy days stayed easy – exactly what they’re supposed to be.
Pace Concerns
Bedrock flagged some inconsistencies across quality sessions. Four key tempo efforts sat in a tight 10-second pace band (4:46/km-4:56/km), but their average heart rates ranged from 161-177 BPM. This is a 16 BPM spread for nearly identical effort outputs.
The agent saw that spike as either a sudden loss of aerobic fitness or poor effort control. In reality, that spread was driven mostly by this summer’s heatwaves, during which thermal stress artificially elevated my heart rate.
This highlights a gap in my initial dataset. Because FIT files don’t record ambient weather, Bedrock was analysing telemetry with insufficient context. Ingesting weather API data into the pipeline could resolve this in future analysis.
Volume Progression
Next, volume progression: the increasing weekly mileage a training plan builds toward race day. This is the main way to build endurance…and also a frequent cause of injury! Push the weekly total up too slowly, and marathon day will arrive without a sufficient endurance base. Push it too fast, and the body doesn’t have time to adapt to the added load.
Volume Findings
The Bedrock agent reported that my Garmin dataset shows textbook periodisation. A gradual climb with deliberate peaks and recovery weeks placed where they were needed. The buildup from 16.94 km in the week of 1 February to 37.63 km in the week of 27 April took 12 weeks, which is an acceptable progression.
Recovery weeks were also well-scheduled. The week of 4 May dropped to 27.8 km (-26%) straight after the 37.63 km peak, and the week of 1 June dropped again to 25.35 km (-40%) after a second peak of 41.93 km. That second recovery week came right before one of the training block’s biggest pushes: 47.13 km in the week of 8 June, a 12.4% rise from the 25 May peak.
Volume Concerns
For volume progression, a commonly cited guardrail is the 10% rule: don’t increase weekly mileage by more than about 10% from one week to the next. The analysis found that two weeks broke this rule.
The week of 27 April saw a 66% jump from 22.66 km to 37.63 km. The week of 8 June went further, up 86% from the previous week’s 25.35 km to 47.13 km. Both are the kind of single-week spike that usually precedes an injury. The recovery weeks before each spike mitigated some of the risk, but it’s not something to make a habit of!
Race Prediction
Bedrock modelled my estimated marathon time three separate ways:
Method 1: Training Pace Analysis
This method works backwards from recent hard efforts. Firstly, identify the pace I can hold at lactate threshold. Then apply the well-established rule that marathon pace is roughly 8-10% slower than threshold pace, since race effort must be sustained for hours rather than minutes.
Recent quality sessions:
- 30 April: 8.18 km at 4:54/km, 170 BPM (threshold)
- 17 May: 22.54 km at 5:06/km, 169 BPM (marathon pace simulation)
- 11 July: 9.84 km at 4:56/km, 177 BPM (near-threshold)
Bedrock estimate: These paces suggest:
- Threshold pace: ~4:54-4:56/km
- Marathon pace (8-10% slower than threshold): 5:18-5:26/km
- Predicted finish: 3:43-3:49
Method 2: Heart Rate Efficiency Model
This method uses heart rate rather than pace as the anchor: it looks at what percentage of max heart rate can be sustained over a race-length effort, then finds the pace that sits in that same effort zone across the full marathon distance.
I can currently sustain 5:04/km at 170 BPM for 32.22 km. Given:
- Max HR appears to be 193-196 BPM.
- 170 BPM = ~87-88% max HR (solid marathon effort zone).
- Cardiac drift is minimal on long runs.
Conservative marathon pace: 5:20-5:25/km at ~170 BPM.
Predicted finish: 3:45-3:49
Method 3: Jack Daniels VDOT
VDOT is a single number that Jack Daniels’ running formula uses to represent overall running fitness. It combines aerobic capacity (VO2max) with running economy (how efficiently oxygen is used) into one score derived from an actual race performance. That score is then used to predict equivalent times at other distances.
The 22.54 km at 5:06/km (1:54:59) suggests a VDOT of approximately 39-41.
- VDOT 40: Predicted finish = 3:49:45
- VDOT 41: Predicted finish = 3:45:00
Consensus Prediction
All things considered, Bedrock currently estimates my finish time at 3:45-3:49 (5:18-5:26/km), basing this on:
- Peak long run of 32.22 km demonstrating distance readiness.
- Easy runs were reliably held within the Zone 2 range, indicating well-controlled training intensity.
- Recent quality work (11 July at 4:56/km) showing retained speed.
- Total volume of 573 km over 28 weeks is solid marathon preparation.
- Long runs executed at appropriate aerobic effort.
Personally, this feels optimistic. I was targeting closer to four hours, and my Garmin currently estimates a marathon time of 3:57. But this isn’t a short race – minutes of drift in estimates are common in endurance events like the marathon.
Bedrock sensibly includes caveats in its analysis. The race outcome heavily depends on factors including:
- Weather conditions: Heat, rain and wind all significantly impact performance.
- Pacing discipline: Going too fast in the first half will burn energy needed for the second half.
- Nutrition and hydration: Strategy in the week of the event and on the day matters just as much as fuelling during the race itself.
- Taper quality: Tapering too soon or too late can affect marathon readiness.
Summary
This project started as a way to make additional use of my Garmin training data, and became a three-account AWS data pipeline: raw FIT files landing in the Data Lake account, being cleansed and curated in the Lakehouse account and then handed to a Bedrock agent in the AIML account for analysis. Along the way, that agent told me things about my training I hadn’t noticed myself, some encouraging, some…not!
This is very much in the early days of the project, and there’s still plenty to work on. The agent needs more testing and refinement to improve analysis consistency, the code needs a proper review, and Data Lake FIT uploads are still something I do by hand rather than automatically. Future posts will get into all of that.
This project’s brought together a personal hobby and a professional one, and it’s been genuinely satisfying to build. I’ll keep tinkering with it between now and October, race day permitting.
If you’d like to support Pendleside Hospice, you can do so through my JustGiving page.
Like this post? Click the button below for links to contact, socials, projects and sessions:
Thanks for reading ~~^~~












