Categories
Data & Analytics

Building An AWS Garmin Data Pipeline For My Charity Marathon

I’m running the Chester Marathon for charity to support Pendleside Hospice! In this post, I build an AWS data pipeline to analyse my Garmin training data.

Introduction

Outside of my AWS and community work, I have several hobbies – one of which is running. I was a big running fan in the 2010s, and after a few years away I got back into it in 2025. I’ve had a pretty decent 18 months with it, and this summer I decided to check off one of my bucket list items by running a marathon for charity.

A side effect of my marathon training has been an influx of training data. While I already have access to platforms like Garmin Connect and FetchEveryone to analyse this data, neither platform quite gives me the full picture I want. My choices were to either pay Garmin for enhanced features or build an AWS Garmin data pipeline. And I’m not one to turn down a good data project!

Firstly, I’ll cover how the data is generated and collected. Then I’ll walk through the three-account AWS architecture that turns the raw files into data a Bedrock agent can analyse. Finally, I’ll cover what that agent actually found after I pointed it at my training data.

Future posts will examine the backend behind each pipeline stage. This one’s about the shape of the whole pipeline and the race it was built for. Let’s begin by looking at my chosen charity.

The Charity

Pendleside Hospice is a registered charity based in Burnley. Founded in 1988, they provide palliative and end-of-life care for people across Burnley, Pendle and Rossendale. They originally focused on cancer care, but now support people with a wider range of life-limiting conditions, including dementia, respiratory and neurological illnesses.

Pendleside have cared for a good friend of mine. Martine Hamer worked alongside her husband Russell at his Burnley salon. Russell’s been cutting my hair for over two decades, and Martine was the one who mixed the bleach, ran the hairdryers and occasionally talked me into a spray tan. After her short battle with breast cancer, she died in Pendleside’s care in February 2010, aged just 35.

MartineHamer

Pendleside cares for over 1,700 people every year across a 10-bed inpatient unit, a Hospice at Home service, day therapy and bereavement support for families and children. They also run a Meals on Wheels service and a Health, Wellbeing & Rehabilitation programme that offers services such as physiotherapy and counselling.

Operating Pendleside requires over £4.5 million a year. An NHS grant covers around 20% of that, leaving over £3.5 million to be raised each year locally through events, their charity shops and donations.

I’ve fundraised for Pendleside before, doing the 2010 Bupa Manchester 10k for them. But there was always something else in the back of my mind…

The Marathon

I’ve thought about running a marathon for charity several times over the years, but never felt right doing it when I was frequently running a couple of them a year. After all, if I was doing that many, where was the challenge? Having had an extended break and several life changes, 2026 feels like the ideal time to give it a go.

While this is my first charity marathon and my first one for a while, it’s not my first marathon. That honour falls to the wet, windy 2012 Manchester Marathon, which still lives in infamy in the minds of those who took part.

Amazingly, I came back for more after this! In the following years, I ran various marathons, including locally in Manchester and in the capital, through vineyards at the Marathon du Médoc and around a running track at the Groundhog Marathon.

MBNA Chester Marathon 5983006 121918 Finish 600

This came to a stop, like many things, during the 2020 pandemic. Running then took a back seat for a few years until I got back into it in 2025 after joining the Steamhaus Vitality scheme. Fast forward 18 months, and I find myself with an entry for the 2026 Chester Marathon and the opportunity to finally fundraise for Pendleside.

I’ve done Chester twice before, in 2013 and 2019. There are a few reasons for choosing it for my 14th marathon:

  • Chester is a lovely city, and the route is fantastic and well supported.
  • The Active Leisure team is a great event organiser with over a decade of experience.
  • The last running event I did pre-pandemic was the 2020 Chester 10k. So spiritually, this feels like I’m picking up where I left off.

If you’d like to support Pendleside Hospice, you can do so through my JustGiving page.

Data Collection

In this section, I’ll explain where the Garmin data in the AWS pipeline comes from: what captures it, what’s actually captured and the format it arrives in.

Device

My current watch is a Garmin Forerunner 745, a mid-range running and triathlon watch that sits between the Forerunner 645 and the Forerunner 945. It tracks everything I need for marathon training: GPS position, heart rate, pace and cadence. It also has a pulse oximeter for blood oxygen, a barometric altimeter and onboard training load and recovery metrics.

Upon completion, activities are uploaded to Garmin Connect via my phone’s Bluetooth.

Garmin Connect

Garmin Connect is the companion platform for data collected by Garmin devices. Activity, health and training data are automatically uploaded here after device syncing. Garmin then turns this data into various reports, visuals and feedback. These can range from the pace and heart rate zones of individual activities to aggregated training loads and historical mileage.

Garmin Connect hero image 1480x975

Crucially for this project, it also lets me download the data in various formats, including the raw FIT files that everything else in this pipeline depends on.

FIT Files

Flexible and Interoperable Data Transfer (FIT) is the file format Garmin devices use to record and share activity data.

Rather than storing one big block of numbers, a FIT file is built from streams of timestamped messages. These capture the same GPS, heart rate and cadence data as the watch itself, plus other data like power and speed, all streamed as the activity happens rather than saved as a single summary.

It’s compact and self-describing by design, making it ideal for Garmin Connect and other training platforms like Strava, TrainingPeaks and FetchEveryone.

Architecture

This section examines the multi-account architecture of the AWS Garmin FIT data pipeline. I discuss the concept of a multi-account strategy and its benefits, then examine each account’s role in the data pipeline.

Multi-Account

Before getting into what each account actually does, it’s worth explaining why this pipeline is split across multiple accounts in the first place.

Multi-Account Strategy

A multi-account strategy is the practice of splitting workloads, environments and admin boundaries across multiple AWS accounts. Each account has a specific, well-defined role, rather than relying on IAM policies or resource tagging for separation inside a single account.

Where a single account relies on carefully maintained policies to keep teams, systems and environments separate, a multi-account strategy makes that separation structural. Each account becomes its own hard boundary instead of a policy-enforced one.

The strategy typically extends beyond individual workloads to environments, products and business areas, and is usually managed centrally through AWS Organizations.

But why bother? What benefits does a multi-account setup offer?

Multi-Account Benefits

AWS has a white paper on this – Organizing Your AWS Environment Using Multiple Accounts. And while it’s aimed at enterprise customers, many of the benefits it covers apply just as well to my multi-account AWS Garmin data pipeline architecture.

Data Access Controls: I’ve worked with data for many years, and ensuring the right data is in the right place and accessible to the right people is a common challenge. AWS accounts are ideal boundaries to resolve this, offering strict access controls and enforcing one-way data flow at the account level. This means, for example, that a Lakehouse process can never overwrite raw data in the Data Lake, and an AI process can only read Lakehouse data that’s been cleansed and curated.

Blast Radius Limitation: Things break, and things fail. When that happens and everything lives in the same account, it can be disastrous. A multi-account setup has inherent boundaries that limit the scope of adverse events such as misconfigurations, compromised credentials or bad deployments.

Grouping By Purpose: Splitting workloads into separate accounts means I can manage them on their own terms without creating unnecessary or unintended dependencies. This separation also ringfences each account’s use of services, quotas and billing, making scaling up far simpler.

Pipeline Accounts

This section reviews each AWS account of the Garmin data pipeline. First is the Data Lake account, where raw data lands and is stored. Next is the Lakehouse account, which cleanses and curates the data. Finally is the AIML account, which analyses the curated data for insights.

Data Lake Account

This account is a landing zone. Its main job is to capture and store raw data and grant read-only cross-account access to it.

It currently holds two S3 buckets: one for sensitive data and a general-purpose one for everything else. Data here is version-controlled, immutable and easy to locate and back up. Data is currently added manually, but over time this will expand to include ingestion scripts that pull data from third parties, as well as automated pushes from agents and IoT sensors.

Lakehouse Account

This account transforms and serves data. It owns the ELT pipeline, the transformed data storage, the Glue Data Catalog and the Athena workgroup. The Lakehouse has Cleansed, Curated and Semantic layers, with further separation for sensitive and general data classifications.

While the Lakehouse account can access both Data Lake buckets, access is scoped per process rather than granted fully. The AWS Garmin data pipeline’s ELT processes can only access the general bucket, as they don’t need anything from the sensitive one. The AIML account’s access works on the same principle of least privilege, using a separate cross-account role scoped to a single Glue database, S3 prefix and Athena workgroup.

AIML Account

This account consumes data from the Lakehouse account for GenAI processes. It owns a Bedrock agent, two Lambda functions and its own S3 buckets for query results and generated output.

The AIML account queries data in the Lakehouse account through its Glue Data Catalog. It has no access to the Data Lake account and cannot read data that has not already passed through the Cleansed and Curated Lakehouse layers. This ensures the agent uses only validated, trusted data.

Processes

This section covers what actually happens at each stage of my AWS Garmin data pipeline: the triggers, the hand-offs and the grants the accounts use to interact with each other.

There are three AWS accounts in the Garmin data pipeline, each of which owns separate processes:

CharityBlog AWSAccounts.drawio
  • Data landing and storage in the Data Lake account.
  • Data cleansing and curation in the Lakehouse account.
  • Analytics and report generation in the AIML account.

Data Lake Account

data lake storage account architecture dark

Garmin data as FIT files enters the raw-general S3 bucket by manual upload. The account then answers read requests from the Lakehouse’s Cleansed ELT Lambda, granted via bucket policy. There is no processing, no deletion and no outbound calls initiated from this account.

Lakehouse Account

lakehouse account architecture

The Lakehouse account owns two distinct processes:

  • The Cleansed process, which converts the raw FIT files into readable Parquet data.
  • The Curated process, which validates the data and derives new columns to add value.

Cleansed Process

garmin cleansed pipeline stepfunctions graph

An EventBridge Scheduler runs a cron trigger every day at 06:00 UTC, which starts a Step Function. In the first step of this execution, an S3 List step reads the contents of the Garmin folder in the Data Lake account’s raw-general S3 bucket.

Next, the Cleansed Filter Lambda checks each FIT object against those already processed and passes along the unprocessed ones. Then a Map state fans out across these objects, running the Cleansed ELT Lambda once per object with a maximum concurrency of ten.

Each ELT invocation pulls its FIT file directly from the Data Lake account, using bucket-policy-granted access on the Lambda’s execution role. It then decompresses the file and parses the FIT binary with garmin-fit-sdk.

From there, it converts FIT epoch timestamps into standard datetimes and semicircle coordinates into decimal-degree latitude and longitude, then writes the results as Parquet files to the Cleansed bucket. 21 tables are written per activity, with partitions and table schemas registered or updated in the Glue Data Catalog.

Curated Process

garmin curated pipeline stepfunctions graph

When the Cleansed Step Functions execution succeeds, an EventBridge Rule triggers a second state machine for the Curate process.

In the first step of this execution, a Curated Filter Lambda identifies activities that have not yet been promoted by checking whether curated output already exists. Then a Choice state checks if everything is up to date. If it is, no further action is taken. If not, a Map state fans out across the unprocessed files, invoking the Curated Promote Lambda for each activity with a maximum concurrency of ten.

Each Curated Promote Lambda invocation reads its assigned Cleansed Parquet file for all 21 tables and computes derived features on the records table: elapsed seconds, cumulative distance, grade percentage, pace (min/km) and heart-rate zone. It then validates every table row against a corresponding Pydantic schema and writes valid rows to the Curated bucket. The Curated tables are registered in a separate Glue database.

Reliability & Access

Idempotency is handled at multiple points in both processes:

  • The Cleansed Filter Lambda won’t include a file in the Map state if its output already exists.
  • Cleansed ELT Lambda invocations perform their own HeadObject check before writing, in case two invocations race each other.
  • Curated Filter and Promote Lambdas apply the same output-exists pattern against the Curated bucket.

Additionally, CloudWatch alarms trigger if any Lambda invocation errors out.

Once the Curated tables are registered, the AIML account can query them via Athena. A cross-account GarminDataCrossAccountReadRole IAM role is assumable by the AIML account’s Garmin Agent Lambda, and grants the following:

  • Read access to the Curated Glue database.
  • Read access to the S3 prefix where those tables live.
  • Use of a dedicated garmin-agent-queries Athena workgroup with an enforced output location.

Results land in a separate athena-results S3 bucket in the Lakehouse account, kept apart from the process’s own working buckets.

AIML Account

aiml account garmin agent architecture

Two separate flows run here. In the interactive flow, a user prompt is sent to the marathon-training-analyst Bedrock Agent, running on Claude Sonnet. The agent interprets it and calls one of six endpoints in the query Lambda’s action group depending on the prompt, covering topics such as training summaries, activity details, heart rate trends and pace trends.

The Query Lambda then assumes GarminDataCrossAccountReadRole in the Lakehouse account via AWS STS and runs an Athena query against the Curated Glue database, scoped to the garmin-agent-queries Athena workgroup. It converts the raw metrics (m/s to min/km pace, metres to km, seconds to H:MM:SS) and returns structured results to the agent, which formats a natural-language response.

In the batch flow, the Orchestrator Lambda calls the same agent with a comprehensive analysis prompt that explores topics such as volume progression, pace trends and race predictions. It collects the agent’s streamed multi-step reasoning and query results, combines them into a Markdown document, and writes the finished output to the genai-results bucket where it’s picked up for personal use.

AIML Agent Prompt

For completeness, I wanted to briefly discuss the marathon-training-analyst agent’s prompt. I won’t include the full prompt here for two reasons:

  • It’s been superseded by a newer version, which I’ll discuss later.
  • The prompt is over 220 words long, so focused sections will be easier to follow than the whole.

So here is an abridged version of what the agent gets when the batch flow runs. I’ve removed some of each section’s bullets – the remaining text is unaltered:

“You are an expert marathon running coach and sports data analyst. You have access to a runner’s Garmin training data and your job is to analyse their marathon preparation.

When analysing training data:

  • Calculate and interpret weekly mileage progression, noting appropriate buildup and recovery weeks
  • Identify injury risk factors: sudden volume spikes (>10% week-over-week), insufficient rest days
  • Provide race day predictions using established methods (Jack Daniels formula, pace-based estimates)

When producing output:

  • Lead with key insights, support with specific numbers from the data
  • Use markdown formatting with headers, bullet points, and emphasis
  • Include specific dates, distances (in km), paces (min/km), and heart rates where relevant

Always query the data before making claims. Do not hallucinate metrics — use only what the data shows.
Garmin speeds are in m/s — convert to min/km for display (pace = 1000/60 / speed_m_s).
Distances in the data are in metres — convert to kilometres for display.”

Isn’t This All Overkill?

A fair question to end on. And yes – this would be total overkill if all I was planning to do with this was the Garmin analysis.

But each of these accounts has more work ahead of it. There will be new ingestion processes in the Data Lake account, new databases in the Lakehouse account and new agents in the AIML account. Indeed, this exact framework will now be receiving Project Wolfie data in the coming months!

The intention behind this architecture is to build something I can use long-term. While it may seem like a lot of work and resources for such a small project, the more time I spend on it now, the more likely it is to support my future requirements.

And that’s the AWS Garmin data pipeline! Three accounts and a lot of careful plumbing, taking raw FIT files and transforming them into queryable data. So what did those 28 weeks of training data actually say?

Tests & Insights

In this section, I’ll go through what the Bedrock agent found when it analysed my Curated dataset generated by my AWS Garmin data pipeline, covering 28 weeks and 66 running activities totalling 573 km.

A brief caveat. These findings are from a version of the agent that has since had some tweaks and changes. These findings remain valid, and I’ll explore the agent’s changes in future posts.

Let’s start with pace trends. This is how pace and heart rate interact across different types of runs. This is important because pace alone isn’t the whole story.

Running at a certain speed with a low heart rate is a very different signal from running at the same speed with a high one, and comparing the two over time is one of the clearest ways to see whether training effort is being optimally controlled. Two of the most suitable run types for this analysis are easy runs and quality sessions.

Easy runs are deliberately slow, low-effort miles that build an aerobic engine without piling on fatigue. The goal here isn’t speed – it’s staying comfortable enough to hold a conversation. The heart rate range that keeps a run ‘easy’ is often called Zone 2, calculated in my case via Heart Rate Reserve rather than straight % max, at roughly 60-70% of HRR. For me, this currently falls somewhere between 145 and 165 BPM.

Quality sessions are structured, harder efforts such as tempo runs or threshold work. These target a specific pace range to build speed and the ability to sustain effort under fatigue.

Pace Findings

The Bedrock agent reported that the easy runs were suitably paced. The bulk of the shorter runs fell in the 5:15-5:35/km range at 145-165 BPM, comfortably within the Zone 2 window. For example:

  • On 17 January, I ran 6.85 km at 5:19 pace and 147 BPM.
  • 08 May: 5.7 km at 5:33 pace and 149 BPM.
  • 17 June: 5.6 km at 5:31 pace and 147 BPM.

Across the 28 weeks, easy days stayed easy – exactly what they’re supposed to be.

Pace Concerns

Bedrock flagged some inconsistencies across quality sessions. Four key tempo efforts sat in a tight 10-second pace band (4:46/km-4:56/km), but their average heart rates ranged from 161-177 BPM. This is a 16 BPM spread for nearly identical effort outputs.

The agent saw that spike as either a sudden loss of aerobic fitness or poor effort control. In reality, that spread was driven mostly by this summer’s heatwaves, during which thermal stress artificially elevated my heart rate.

This highlights a gap in my initial dataset. Because FIT files don’t record ambient weather, Bedrock was analysing telemetry with insufficient context. Ingesting weather API data into the pipeline could resolve this in future analysis.

Volume Progression

Next, volume progression: the increasing weekly mileage a training plan builds toward race day. This is the main way to build endurance…and also a frequent cause of injury! Push the weekly total up too slowly, and marathon day will arrive without a sufficient endurance base. Push it too fast, and the body doesn’t have time to adapt to the added load.

Volume Findings

The Bedrock agent reported that my Garmin dataset shows textbook periodisation. A gradual climb with deliberate peaks and recovery weeks placed where they were needed. The buildup from 16.94 km in the week of 1 February to 37.63 km in the week of 27 April took 12 weeks, which is an acceptable progression.

Recovery weeks were also well-scheduled. The week of 4 May dropped to 27.8 km (-26%) straight after the 37.63 km peak, and the week of 1 June dropped again to 25.35 km (-40%) after a second peak of 41.93 km. That second recovery week came right before one of the training block’s biggest pushes: 47.13 km in the week of 8 June, a 12.4% rise from the 25 May peak.

Volume Concerns

For volume progression, a commonly cited guardrail is the 10% rule: don’t increase weekly mileage by more than about 10% from one week to the next. The analysis found that two weeks broke this rule.

The week of 27 April saw a 66% jump from 22.66 km to 37.63 km. The week of 8 June went further, up 86% from the previous week’s 25.35 km to 47.13 km. Both are the kind of single-week spike that usually precedes an injury. The recovery weeks before each spike mitigated some of the risk, but it’s not something to make a habit of!

Race Prediction

Bedrock modelled my estimated marathon time three separate ways:

Method 1: Training Pace Analysis

This method works backwards from recent hard efforts. Firstly, identify the pace I can hold at lactate threshold. Then apply the well-established rule that marathon pace is roughly 8-10% slower than threshold pace, since race effort must be sustained for hours rather than minutes.

Recent quality sessions:

  • 30 April: 8.18 km at 4:54/km, 170 BPM (threshold)
  • 17 May: 22.54 km at 5:06/km, 169 BPM (marathon pace simulation)
  • 11 July: 9.84 km at 4:56/km, 177 BPM (near-threshold)

Bedrock estimate: These paces suggest:

  • Threshold pace: ~4:54-4:56/km
  • Marathon pace (8-10% slower than threshold): 5:18-5:26/km
  • Predicted finish: 3:43-3:49

Method 2: Heart Rate Efficiency Model

This method uses heart rate rather than pace as the anchor: it looks at what percentage of max heart rate can be sustained over a race-length effort, then finds the pace that sits in that same effort zone across the full marathon distance.

I can currently sustain 5:04/km at 170 BPM for 32.22 km. Given:

  • Max HR appears to be 193-196 BPM.
  • 170 BPM = ~87-88% max HR (solid marathon effort zone).
  • Cardiac drift is minimal on long runs.

Conservative marathon pace: 5:20-5:25/km at ~170 BPM.

Predicted finish: 3:45-3:49

Method 3: Jack Daniels VDOT

VDOT is a single number that Jack Daniels’ running formula uses to represent overall running fitness. It combines aerobic capacity (VO2max) with running economy (how efficiently oxygen is used) into one score derived from an actual race performance. That score is then used to predict equivalent times at other distances.

The 22.54 km at 5:06/km (1:54:59) suggests a VDOT of approximately 39-41.

  • VDOT 40: Predicted finish = 3:49:45
  • VDOT 41: Predicted finish = 3:45:00

Consensus Prediction

All things considered, Bedrock currently estimates my finish time at 3:45-3:49 (5:18-5:26/km), basing this on:

  • Peak long run of 32.22 km demonstrating distance readiness.
  • Easy runs were reliably held within the Zone 2 range, indicating well-controlled training intensity.
  • Recent quality work (11 July at 4:56/km) showing retained speed.
  • Total volume of 573 km over 28 weeks is solid marathon preparation.
  • Long runs executed at appropriate aerobic effort.

Personally, this feels optimistic. I was targeting closer to four hours, and my Garmin currently estimates a marathon time of 3:57. But this isn’t a short race – minutes of drift in estimates are common in endurance events like the marathon.

Bedrock sensibly includes caveats in its analysis. The race outcome heavily depends on factors including:

  • Weather conditions: Heat, rain and wind all significantly impact performance.
  • Pacing discipline: Going too fast in the first half will burn energy needed for the second half.
  • Nutrition and hydration: Strategy in the week of the event and on the day matters just as much as fuelling during the race itself.
  • Taper quality: Tapering too soon or too late can affect marathon readiness.

Summary

This project started as a way to make additional use of my Garmin training data, and became a three-account AWS data pipeline: raw FIT files landing in the Data Lake account, being cleansed and curated in the Lakehouse account and then handed to a Bedrock agent in the AIML account for analysis. Along the way, that agent told me things about my training I hadn’t noticed myself, some encouraging, some…not!

This is very much in the early days of the project, and there’s still plenty to work on. The agent needs more testing and refinement to improve analysis consistency, the code needs a proper review, and Data Lake FIT uploads are still something I do by hand rather than automatically. Future posts will get into all of that.

This project’s brought together a personal hobby and a professional one, and it’s been genuinely satisfying to build. I’ll keep tinkering with it between now and October, race day permitting.

If you’d like to support Pendleside Hospice, you can do so through my JustGiving page.

Like this post? Click the button below for links to contact, socials, projects and sessions:

SharkLinkButton 1

Thanks for reading ~~^~~

Categories
Data & Analytics

Practical Lakehouse Architecture By Gaurav Thalpati

In this post, I review Gaurav Ashok Thalpati’s 2024 book ‘Practical Lakehouse Architecture‘ published by O’Reilly Media.

Introduction

I first found O’Reilly books a few years back in a Data Engineering-themed Humble Bundle. Since then, I’ve built an extensive library of both e-books and physical books, with many more on my Amazon wish list. At the start of 2025, I decided to actually start reading them…

So far, I’ve finished three. Now, I don’t feel compelled to review them all. But having finished Practical Lakehouse Architecture I decided to start the Shark Shelf. This will be an occasional series of review posts about books that I really like, or that deserve some fanfare. And yes – How To Solve It belongs on the Shark Shelf.

Now let’s talk about Practical Lakehouse Architecture.

The Author

Gaurav Ashok Thalpati hails from Pune, India, where he’s worked as an independent cloud data consultant for decades. He’s a blogger and YouTuber, holds multiple data certifications and is an AWS Community Builder.

In July 2024, O’Reilly published his first book, Practical Lakehouse Architecture.

The Book

From the Practical Lakehouse Architecture blurb:

This guide explains how to adopt a data lakehouse architecture to implement modern data platforms. It reviews the design considerations, challenges, and best practices for implementing a lakehouse and provides key insights into the ways that using a lakehouse can impact your data platform, from managing structured and unstructured data and supporting BI and AI/ML use cases to enabling more rigorous data governance and security measures.

Practical Lakehouse Architecture was released in July 2024. It is available in both physical and eBook forms from O’Reilly, Amazon US, Amazon UK and eBooks.

Motivations

Reading a book?! In 2025?! I know, right? This section examines my motivations for buying and reading Practical Lakehouse Architecture.

Project Wolfie

I recently wrote about the beginning of Project Wolfie. I kinda expected to have started coding by now. Instead, most of my work is currently on paper and whiteboards. But there’s a good reason for this.

Project Wolfie is greenfield. I don’t have any existing code or resources, and I can use modern tools freely. However, with this freedom comes responsibility. Every choice I make now affects the architecture and involves tradeoffs. As much as I want to start working on the deliverables, I also want to make sensible decisions that can withstand scrutiny.

My hope with Practical Lakehouse Architecture was that it would help me with critical areas like observability, CI/CD, and security. Because it’s not that there isn’t advice online…

Advice Spread Thin

Lakehouse architectures are relatively recent in the data landscape. As a result, their understanding is not as established as that of data warehouses and data lakes, and some aspects of Lakehouse architecture are still evolving.

Many Lakehouse resources are either brief overviews, opinionated deep dives into specific use cases or marketing posts acting as best practices. This makes it hard to find balanced advice. My hope with Practical Lakehouse Architecture was that it would offer clear, unbiased views.

Professional Curiosity

As of 2025, I’ve spent nearly a decade in technical data roles. And in that time I’ve seen massive changes in data management, ranging from a server cupboard in Stockport to huge, multi‑region distributed data platforms.

Over the years, I’ve cultivated a passion for data technology, evolving from writing blog posts and speaking at meetups to working as an AWS consultant. As an AWS Community Builder in the Data category, I can access early previews and best practices from AWS experts. Additionally, as an AWS User Group Leader, I help attendees and guest speakers discuss data patterns.

With this in mind, I was curious about what new insights Practical Lakehouse Architecture could offer me.

Book Review

Onto the review! In this section, I’ll summarise the chapters and examine what stood out in each.

Chapters 1 – 3

The first set of chapters introduces the foundations of Lakehouse architecture, comparing it with traditional models and exploring the importance of storage in modern data platforms.

Chapter 1: Introduction to Lakehouse Architecture lays the groundwork for the book, putting all readers on equal footing for the chapters ahead. Gaurav starts by defining and exploring the ideas and concepts of various data architectures. He then examines the characteristics, evolution and benefits of the Lakehouse architecture.

Chapter 1 can be viewed on the O’Reilly site.

Chapter 2: Traditional Architectures and Modern Platforms contrasts the Lakehouse architecture with traditional data lakes and data warehouses, outlining the benefits and limitations of each. Gaurav then shifts his focus to how modern cloud platforms have transformed these traditional architectures.

I like how Gaurav hasn’t dismissed lakes and warehouses here. Both are proven and well-understood options, and they are still the better choice in certain situations over Lakehouses.

Chapter 3: Storage: The Heart Of The Lakehouse examines the various factors surrounding data storage. Gaurav looks at row-based and column-based storage formats. He then explains the features and uses of Parquet, ORC, and Avro. He also compares newer open table formats, like Iceberg, Hudi, and Delta Lake, highlighting their similarities, differences, and use cases.

This is one area where the book really shines. Having topics like this explained clearly in one place, without having to go online, is incredibly useful!

Chapters 4 – 6

Next, these chapters focus on the operational and organisational elements of Lakehouse architectures. Topics include metadata management, compute engines, and governance. These elements are essential for effectively scaling and securing a modern data platform.

Chapter 4: Data Catalogs explores the purpose of data catalogs and the different types of metadata they can contain. It explains how catalogs support essential processes such as classification, governance, and lineage. Gaurav also compares data catalog implementations across AWS, Azure, and GCP.

Including multi-cloud examples both broadens the chapter’s scope and reinforces the cloud-agnostic nature of Lakehouse architecture – an important theme of the book.

Chapter 5: Compute Engines for Lakehouse Architectures examines compute options for batch and real-time data processing. Gaurav covers open-source tools such as Spark, Flink, and Presto, as well as cloud-native services like AWS Glue, Google BigQuery, and Databricks. He offers practical advice for selecting a compute engine, considering factors such as provisioning complexity, open-source support and AI/ML capabilities.

Chapter 6: Data and AI Governance and Security in Lakehouse Architecture explores governance and security, crucial areas for any production-ready data platform. Gaurav discusses core topics such as data quality, ownership, sensitivity and compliance. He also explores how governance responsibilities span both business and technical domains, emphasising the importance of organisational roles in maintaining control and oversight.

Chapters 7 – 9

Finally, these chapters focus on the practical realities of Lakehouse implementation – moving between theory and practice, and looking ahead to the architecture’s potential future.

Chapter 7: The Big Picture: Designing and Implementing a Lakehouse Platform examines considerations ranging from requirements gathering to defining business goals. Recommended Lakehouse zones are analysed and explained, and the expectations for each zone are defined. Finally, CICD is considered, and a sample design questionnaire is provided to help guide implementation planning.

Zones, or layers, are currently one of the most contentious areas of Lakehouse architectures. I like Gaurav’s stance on this – it’s somewhat similar to Simon Whiteley‘s. Yup – this video again.

Chapter 8: Lakehouse in the Real World does something I don’t see often – contrasting ideal scenarios with real-world events. It covers key stages in a Lakehouse’s development like analysis, testing and maintenance, examining what could go wrong and offering mitigation strategies.

This section is definitely accurate, as I’ve encountered some of these factors! It includes comparing greenfield and brownfield implementations, examining how business constraints affect technology choices, and considering if the desired RPO and RTO targets are financially and logistically possible.

Finally, Chapter 9: Lakehouse Of The Future looks ahead, exploring how Lakehouses might evolve in the years to come. Gaurav discusses potential intersections with trends like Data Mesh, Zero ETL and AI model integration. He also introduces emerging technologies like Delta UniForm and Apache XTable, which aim to improve interoperability across data processing systems and query engines. Finally, he touches on future innovations such as Apache Puffin and Ververica Streamhause that could further transform the data landscape.

(Sidenote: this Dremio post explores UniFrom and XTable very well.)

Thoughts

Having finished the book (in two weeks no less!), here are my thoughts:

Firstly, it’s not an intimidating read. At 283 pages, Practical Lakehouse Architecture is authoritative and content-rich without being overly complex or wordy. It also uses familiar O’Reilly conventions and style. When placed next to similar books I own, like The Data Warehouse Toolkit (600 pages) and Designing Data-Intensive Applications (614 pages), it’s easier to pick up and get into. And with some books, that’s a battle in itself!

PXL 20250417 143214247~2

Also, Practical Lakehouse Architecture‘s flow is very natural and the chapters make their points very well. I find some technical books, including some O’Reilly ones, hard to follow because they feel disjointed and jargon-heavy. That wasn’t the case here. The book held my attention very well throughout, and will serve me well as a future reference point.

Practical Lakehouse Architecture also feels like it will be relevant for a while. Some of my technical books have sections that are now outdated due to rapid technological changes. Here, ideas such as decoupled storage and compute, unified governance, and data personas will continue to matter for years to come.

Overall, an excellent book that I enjoyed reading.

Summary

In this post, I reviewed Gaurav Ashok Thalpati’s 2024 book ‘Practical Lakehouse Architecture‘ published by O’Reilly Media.

Ultimately, Practical Lakehouse Architecture is a well-written and informative book that caters to a wide range of skills. It’s a strong addition to the O’Reilly catalogue and complements titles like Rukmani Gopalan‘s 2022 book, The Cloud Data Lake, which I’m currently reading. It’s a great knowledge source for this constantly evolving modern data architecture.

If this post has been useful then the button below has links for contact, socials, projects and sessions:

SharkLinkButton 1

Thanks for reading ~~^~~

Categories
Me

Project Wolfie: Honouring An Absent Friend With Music

In this post, I discuss our late German Shepherd Wolfie and outline a project that utilises music data in his memory.

Introduction

On 24 January 2025, our German Shepherd Wolfie sadly passed away.

PXL 20230501 121501481 min

Wolfie was a long, fluffy boi with a big presence and a big mane. He enjoyed sniffing things, tilting his head and barking at foxes. He was loud, proud and bushy-browed.

Wolfie was also psychic. No one could leave the house without his knowledge, and no one could enter it without having a snout-first search.

IMG 20200923 121128 min

Wolfie struggled with various health issues throughout his life, including eating difficulties, muscle problems and genetic defects. Despite this, Wolfie was a cherished family member for four years before he sadly lost his battle with kidney failure and suspected cancer.

Our walks often involved music, and I regularly exposed Wolfie to my music library. He grew so accustomed to iPods during walks that he would bark whenever I picked one up!

PXL 20220904 125315538.MP min

After he was gone, I found myself revisiting the songs we shared. Music became a way to cherish those memories, so I wanted to create something meaningful in his memory…

Project Wolfie

This section explores what Project Wolfie is, the music data it utilises and its goals.

Definition

This project has been on my mind for some time now, and I suppose this was the push it needed to take shape. Project Wolfie is a data-driven initiative that explores the patterns hidden in my music collection. It analyses track metadata, listening habits and technical attributes to find insights, trends and recommendations.

Here, Wolfie is short for:

Waveform Observations Library For Intelligence Engineering

Let’s break this down:

Waveform: A visual illustrating a track’s traits like timbre, pitch and dynamics. Time is represented on the horizontal axis, while the vertical axis reflects amplitude.

Here is a sample waveform:

Observations Library: A consolidated data repository containing information about my music’s properties and my listening habits. The data consists of various types, structures and formats, and will be stored, cleaned and enriched for further use.

For Intelligence Engineering: The AI and BI use cases for the observations library. Here, interactive data visualisation and machine learning services will use the data to uncover patterns, predict trends and generate personalised recommendations.

Data

Music files contain more than just sound – they hold layers of metadata that are crucial to Project Wolfie.

This section explores the different types of metadata related to my music collection, highlighting their functions and purposes. I have assigned these categories using my understanding and intended use of the data.

Technical Metadata

Technical Metadata refers to the measurable and technical attributes of a music file. It tends to include numerical values and audio properties, and is commonly found by analysing the track using applications like Audacity, foobar2000 and MixedInKey, as well as Python libraries like Librosa.

Examples include:

  • What is the track’s initial tempo and key?
  • What is the track’s duration, and how loud is it?
  • What are the track’s spectrographic and harmonic properties?

Descriptive Metadata

Descriptive Metadata refers to the contextual and identifying information about a music track. It tends to include text-based details and is commonly found both within the track’s properties and on websites like Beatport and Discogs.

Examples include:

  • Who produced the track, and what is it called?
  • What is the track’s genre?
  • Which label published the track, and when?

Interaction MetaData

Interaction Metadata refers to engagement and listening behaviours. It typically includes dates, integers and timestamps, and is commonly generated by digital music services like iTunes and Spotify.

Examples include:

  • When was the last time a track was played or skipped?
  • How many times has a track been played?
  • What rating has a track been assigned?

Deliverables

Here are the objectives I’m pursuing in Project Wolfie. Given their complexity, they will be divided into multiple epics and spread out over an extended period.

Data Lakehouse

So far, I have discussed the importance, types, and applications of data. To this end, I need to fulfil a few requirements:

  • Ingesting and storing data from multiple sources.
  • Transforming and cleaning data at scale.
  • Enriching and aggregating data for analytics and consumption.

In short, I need a Data Lakehouse. I’ve written about them before and have followed the Medallion Architecture through bronze, silver and gold layers. For Project Wolfie and moving forward, I’ll be using the well-documented and supported AWS reference architecture:

I find this clearer and more regimented than the Medallion Architecture. It also aligns with the points made in Simon Whiteley‘s Advancing Analytics video, which I agree with.

Of course, a good Data Lakehouse isn’t possible without good data…

Quality & Observability

A Data Lakehouse’s effectiveness depends on data quality and observability. Project Wolfie must address factors like:

Veracity & Validation Checks: Verify data accuracy. Checks such as schema validation, null checks and data quality rules can identify issues early, stopping incorrect data from propagating downstream.

Anomaly Detection: Identify patterns often missed by validation like volume spikes and missing periods. Timely anomaly detection shields downstream resources from requiring remedial measures and lowers unforeseen cloud and developer expenses.

Lineage Tracking: Track the data’s journey from ingestion to consumption, documenting all transformations and processes. Vital for debugging, auditing and validation.

Governance & Security

A Data Lakehouse must balance accessibility and control. Governance and security protocols protect data while encouraging responsible usage.

I own all Project Wolfie data, so I have permission to process it. Additionally, there is no sensitive information or PII. However, there are other factors to consider:

Access Controls: Establish guidelines for who and what can access Project Wolfie resources. This safeguards data and services from unauthorised access, misuse and malicious activities.

Data Controls: Establish criteria for availability, backups, and structure. This aids in managing costs, ensuring disaster recovery, and maintaining schema consistency.

Monitoring & Logging: Track access patterns and record changes to data and infrastructure. This improves visibility into both potential threats and cost-related opportunities and vulnerabilities.

AI & BI Use Cases

Finally, I want to extract value and insights from Project Wolfie using Artificial Intelligence (AI) and Business Intelligence (BI). I have data from 2021 onwards from a music collection I started in the early 2000s, so I have lots to work with!

BI Use Cases (Dashboards, Analytics, Insights)

Listening Trends: Identify traits of my collection’s most frequently played and best-represented music. Analyse listening patterns over time to find trends.

Library Optimisation: Find rarely played tracks to add to playlists. Recognise songs that are often played and recommend alternatives for variety.

Distribution Analysis: Analyse my collection’s main genres, publishers and record labels, and investigate the connections between different elements (e.g., “The most popular tracks are typically in the 120-130 BPM range”). Create reports that show diversity and spread (e.g., “90% of house tracks are in five minor keys”).

AI Use Cases (Machine Learning, Automation, Predictions)

AI-Powered Personalised Playlists: Create playlists using the existing library based on properties like BPMs, keys and previous listening patterns, similar to Spotify Wrapped.

Smart Music Recommendations: Use collaborative filtering to suggest search criteria for new music based on my existing collection and listening habits (e.g., “Try G minor tracks at 128 BPM from the early 2010s”).

Predictive Analysis: Use Technical and Descriptive Metadata from new tracks to predict how they will be rated based on my existing library’s metadata (e.g., “This track has harmonic similarities to 70% of your highly rated tracks.“).

Summary

In this post, I discussed our late German Shepherd Wolfie and outlined a project that utilises music data in his memory.

Wolfie enjoyed scent games and retrieving toys, making Project Wolfie’s mission to find and return data and insights a fitting tribute. As the project evolves, I will strengthen its capabilities using new architectures and technologies, honouring Wolfie’s spirit – one track at a time.

Wolfie was more than just a pet; he was a companion and a guardian each day. I miss you big man. Take care out there.

PXL 20211218 180730545.MP min bw

If this post has been useful then the button below has links for contact, socials, projects and sessions:

SharkLinkButton 1

Thanks for reading ~~^~~