Categories
DevOps & Infrastructure

Resurrecting A 2011 PC Into A Kubernetes CKA Lab

In this post, I document the journey of resurrecting a dusty old 2011 PC into a functional Kubernetes lab to prepare for my CKA exam.

Introduction

This year, I plan to take the Certified Kubernetes Administrator (CKA) exam. I’m familiar with tech certifications, having earned my first one in 2019. However, the CKA differs from these exams in that it is a performance-based test that requires completing multiple tasks from the command line with Kubernetes.

I’ve been using Mumshad Mannambeth’s Udemy course, and while it includes several labs I also want to get some extra practice. So I started looking into setting up a Kubernetes lab, and found that I already had access to several devices capable of running one! I have since successfully created my lab, and this post examines how I got there.

I’ll begin by explaining my decision to use the 2011 PC for my Kubernetes lab over other available options, as well as the repairs it needed to get going. Next, I’ll describe the process I followed to select a suitable operating system. Finally, I’ll outline the steps taken to test the PC and install the necessary lab components.

Hardware Selection

In this section, I explain why I chose the 2011 PC for my Kubernetes lab over other available options. While this may seem redundant given the blog title, I wanted to share my considerations and thought processes.

I considered three options:

  • Raspberry Pi Cluster
  • Amazon EC2 Instances
  • A 2011 PC

Raspberry Pi Cluster

The Raspberry Pi is a compact single-board computer widely used for building low-power projects and clusters. It enables users to run Linux and explore different architectures, providing an efficient way to learn about distributed systems without the energy costs of traditional servers.

While the humble Pi may not be the most obvious choice for a Kubernetes lab, a Kubernetes Up & Running appendix covers building a Kubernetes cluster on Raspberry Pi. So it was worth considering!

Raspberry Pi Cluster Pros

  • No Virtualisation Overhead: Running Kubernetes on bare-metal Raspberry Pis removes the need for a hypervisor, allowing all ARM processor cycles and RAM to be fully dedicated to the lab.
  • True High Availability Testing: A Raspberry Pi cluster uses a “shared-nothing” architecture. Unplugging a node forces the cluster to demonstrate its real-time self-healing capabilities.
  • Energy Efficiency: While an old PC draws 60W–100W, a 3-node Pi cluster runs on under 15W. This makes 24/7 operation significantly cheaper.

Raspberry Pi Cluster Cons

  • High Capital Expense: A starter kit with nodes, reliable power supplies and network gear easily exceeds £100 in 2026. I already own a Pi 4, but I’d still need everything else.
  • Component Rigidity: Unlike a desktop PC with swappable RAM, Pi memory is soldered in place. Increasing a Pi’s RAM requires a complete hardware refresh, which limits the lab’s long-term lifecycle.
  • Storage Fragility: Constant logging in Kubernetes can degrade standard microSD cards within months. Reliable operation usually requires booting from USB SSDs, which increases both cost and physical clutter.

EC2 Instances

Amazon EC2 is a standard service for on-demand virtual servers that lets users rent computing capacity in AWS data centres. It operates on a “pay-as-you-go” model, allowing users to launch and terminate instances with the required CPU and RAM as needed.

Note that Amazon also provides Elastic Kubernetes Service (EKS). However, EKS is a managed service and I want to practise the manual Kubernetes setup process.

EC2 Pros:

  • Production Parity: EC2 offers a realistic environment with the same features and specs as a professional enterprise deployment.
  • On-Demand Scalability: Complex node architectures can be quickly created or deleted, enabling the testing of scenarios that would take hours to set up on physical hardware.
  • No Hardware Maintenance: The cloud provider manages the entire bare-metal abstraction. Users are not responsible for managing the hypervisor, physical cooling or local network switching.

EC2 Cons:

  • Ongoing Costs: Even small t3.micro instances accumulate costs quickly. A three-node lab running 24/7 would eventually cost more than the other options.
  • Reduced Learning Opportunities: AWS manages most of the backend networking. This limits opportunities to learn about topics such as network bridges, VLANs or physical NIC teaming.
  • Internet Dependency: The lab’s reliability depends on my home internet connection. If it drops, the entire cluster is unreachable.

A 2011 PC

The time has come. Meet the HP Pavilion p6-2022uk:

Tower 700
The joys of photographing something with black shiny plastic – Ed

This BEAST originally shipped in 2011 with specs including:

  • Processor: Intel Core i3-2120 (3.3 GHz, Sandy Bridge)
  • Memory: 4 GB DDR3-SDRAM (1333 MHz), 2 slots
  • Storage: 1 TB HDD (typically 7200 RPM)
  • Operating System: Windows 7 Home Premium 64-bit
Internals 700

On the back are four of the PC’s six USB 2.0 ports (none of your fancy USB 3.0 here), three audio ports, VGA and DVI outputs (no HDMI) and the Ethernet port:

Rear 700

So how does it stack up?

HP Pavilion Pros:

  • No Capital Expense: Turning existing hardware into a learning tool prevents e-waste and incurs no cost.
  • Full Hardware Control: The PC provides full access to the hardware BIOS, the physical network stack and the hypervisor kernel. This offers an in-depth understanding of how the lab works.
  • Expandability: The PC allows for modular upgrades, making it easy to add more RAM, replace a hard drive or install new network and graphics cards. In contrast, these upgrades can be challenging or even impossible with a cluster of Raspberry Pis.

HP Pavilion Cons:

  • Power Inefficiency: The PC must stay on while the lab operates. This will use more power than the other options, and the power supply, which is from 2011, may be less efficient than modern ones.
  • Legacy Components: The ageing processor, RAM, and ports mean that using the PC will not offer as smooth an experience as a group of Raspberry Pis or EC2 instances. It will lack speed, portability, and hardware redundancy.
  • Repairs Needed: So this leans into why I have this PC in the first place. It do be poorly!

Its biggest problem is the hard drive, which currently boots to an imminent failure warning before failing entirely:

HardDrive 700

The WiFi card is also unreliable. This wasn’t a major issue in the past because it was positioned next to the router and had a constant Ethernet connection. However, this limited its placement options, and it poses a problem now.

Hardware Decision

So why, with all these negatives, did the PC win? It came down to control, simplicity and cost.

Infrastructure Control

With a single physical device, I have complete control over every layer of the stack. I can customise the storage, upgrade the RAM or change the internals whenever necessary. This is impossible with a Pi cluster and limited with EC2 instances.

Minimal Operational Complexity

A PC simplifies everything by combining all components into one device, eliminating the need to manage multiple setups. This way, I can focus on learning Kubernetes instead of juggling cables and managing security groups.

Cost Efficiency

Using a PC lets me avoid the high capital expenditure costs of a Raspberry Pi cluster or the ongoing operational expenses of EC2. Instead, I can invest in repairs and upgrades for the PC and eventually replace it when it fails.

Next, let’s get onto those repairs!

Repairing The PC

If this 2011 PC is going to run Kubernetes, it needs some repairs! In this section, I’ll triage the PC and attempt to fix its issues. For clarity, those issues are prioritised in the following order:

  • Dead hard drive
  • Low RAM
  • Questionable WiFi card

Hard Drive Replacement

Addressing the biggest problem first, that knackered hard drive must go! While fixing hard drives isn’t impossible, it needs specialised skills and is labour-intensive. And in this case, it’s not worth the effort. There’s nothing valuable on the drive, and due to its age it could fail at any moment. That makes this a replacement job!

Thankfully, I made this decision a while ago – well before the current hard drive cost spike! I realised I needed neither an SSD’s speed nor a terabyte of storage. So I bought a Seagate Barracuda ESA-3502 500 GB HDD as it reviewed well and came in at a bargain price of £19.99 (as opposed to the current £35.99 asking price!)

RAM Upgrade

Next, let’s talk about RAM. Although Kubernetes can operate with just 4GB of memory, the virtual machines and the PC also require some memory to function properly. Having such a low memory limit will lead to various problems. The VMs may become too slow, and the PC itself might encounter performance issues.

Over the years, I’ve owned many computers. And like many people in the IT profession, I have various cables, heat sinks, and parts salvaged from computers that were destined for the scrap heap!

Among these items, I found two 8GB DDR3 RAM sticks which I swapped for the PC’s two 2GB sticks. Now this may not work, as the PC’s specs specify that each slot’s maximum memory capacity is 4GB. No harm in trying though, eh?

Internet Access

Finally, let’s discuss the WiFi situation. The WiFi card is already known to have issues, and the PC is on a different floor from my home router so I can’t use Ethernet. This won’t affect the operating system installation, as both the USB ports and the optical drive work. But I won’t then be able to download any updates, firmware or the various Kubernetes components.

Enter my Raspberry Pi 4! This is fully capable of connecting to my WiFi and sharing that connection with my PC. It acts as a NAT (Network Address Translation) router, managing the communication between its wireless interface (wlan0) and its wired Ethernet interface (eth0). This setup is similar to how a home router shares a single public IP address among multiple devices.

Here is what happens at the network layer:

  • The Pi connects to the WiFi network via wlan0 and receives an IP address from the router (e.g. 192.168.1.12). The router sees the Pi as another wireless client.
  • The Pi assigns itself a static IP on eth0 in a separate subnet (e.g. 10.0.0.1) and runs a small DHCP server on that interface via dnsmasq. The PC will use this network.
  • The Pi and PC’s Ethernet ports are then wired together.
  • When the PC sends traffic intended for the internet, it is routed to the Raspberry Pi, which serves as its default gateway.
  • The Raspberry Pi modifies the packet’s source IP address to match its own wlan0 address before forwarding it upstream. This process is known as NAT Masquerading.
  • Incoming replies are received by the Raspberry Pi, which translates the IP addresses back to their original format and then forwards them to the PC.
PiNAT 600

That’s everything! The PC should now boot. But boot what?

OS Selection

As I’m replacing the PC’s hard drive, I can also upgrade the operating system. Since I plan to run virtual machines on the 2011 PC for my Kubernetes nodes, the host OS will act as a hypervisor.

It will handle three critical tasks:

  1. Resource Management: Allocating CPU and RAM for each VM.
  2. Storage: Managing how the VMs access the physical hard drive.
  3. Networking: Creating a bridge so VMs can talk to each other and the internet.

Keep in mind that the OS selections for the 2011 PC and for the Kubernetes VMs are independent choices and needn’t be the same. For instance, Windows can run WSL, and a Debian host can run Ubuntu VMs. The decision on the VM OS will be made after the host OS has been determined.

Let’s start with the traditional players. Which, it turns out, aren’t all that suitable…

Windows

Windows is the most widely used desktop OS and provides a professional-grade virtualisation platform through Hyper-V. While it is a strong option for running enterprise mixed-OS environments, its architecture is poorly suited for a dedicated Linux-based lab on legacy hardware.

  • Tooling Incompatibility: The standard open-source orchestration stack, including virt-manager, cloud-init, and virsh, is designed for the Linux/KVM ecosystem. Managing a Windows host on the 2011 PC needs a distinct set of PowerShell tools, adding unhelpful complexity to the learning process.
  • Resource Inefficiency: A Windows host uses a significant amount of RAM and CPU cycles to operate its GUI and background telemetry. On a dual-core i3-2120 with 4GB of RAM, this operational overhead leaves less room for the guest VMs.
  • Storage Latency: The NTFS file system, along with Windows’ aggressive background indexing and telemetry, creates significant disk contention. This latency becomes more pronounced on older mechanical drives, leading to slower boot times and reduced responsiveness for virtual disk images.

macOS

macOS is well known for providing an exceptional developer experience and is the primary OS used by many in the DevOps community. However, its kernel is optimised for use as an interactive workstation rather than as a persistent, headless hypervisor host.

  • Kernel Abstraction: macOS does not have a built-in hypervisor compatible with Linux, so tools like Docker Desktop or Lima must run a helper VM to create a Linux environment. This results in nested virtualisation, where one VM runs inside another, complicating networking and reducing performance.
  • Architecture Mismatch: Most production Kubernetes clusters are designed for amd64 architecture. Running amd64 nodes on an Apple Silicon Mac involves using QEMU for emulation. This method reduces performance and can introduce subtle behavioural bugs in low-level networking and storage drivers that don’t appear on native hardware.
  • Hardware Inflexibility: macOS is optimised for interactive power management. When used as a persistent, headless server, the OS often deprioritises background tasks or enters sleep states. This can disrupt the high-availability heartbeats needed by a Kubernetes cluster.

Linux

Linux is the industry standard for virtualisation. KVM (Kernel-based Virtual Machine) is built directly into the Linux kernel, making virtualisation a primary feature rather than an additional component. Linux also spotlights the key components used by both hypervisors and Kubernetes: namespaces, cgroups, and network bridges, without the licensing issues of proprietary operating systems.

Linux has many distributions, each with its own package manager, default software and design philosophy. I considered the following Linux distros:

Alpine Linux

Alpine is designed for minimal-footprint environments where reducing attack surface and efficient resource use is crucial. It serves as the base image for many public Docker containers and is employed in embedded systems and security-focused infrastructure, where any unnecessary package can pose a risk.

Alpine Pros
  • Extreme Resource Efficiency: Alpine generally uses 50-80 MB of RAM while idle. This optimisation reduces host-level overhead on older hardware, maximising resources for guest VMs.
  • Reduced Attack Surface: Alpine minimises potential attack vectors by removing unnecessary binaries and packages.
  • Lightweight Init System: It uses OpenRC instead of systemd, resulting in quicker boot sequences and service management through transparent shell scripts.
Alpine Cons
  • GNU Toolchain Friction: Alpine replaces the standard GNU utilities with BusyBox. However, resources such as the KVM documentation assume GNU, so commands must be translated for BusyBox.
  • Standard Library Divergence: The use of musl libc instead of glibc can lead to silent failures and unexpected problems with precompiled binaries or certain kernel modules.
  • Manual Configuration: In Alpine, certain administrative tasks that are automated in other distributions, such as kernel module persistence and complex network bridging, require manual intervention.

Debian Linux

Debian is designed for stable, long-term server infrastructure, prioritising predictability over having the latest software. It serves as the foundation for several distros, including Ubuntu and Raspberry Pi OS, and has built a reputation based on a careful release process that emphasises reliability.

Debian Pros
  • Predictable Performance: A minimal Debian installation uses less than 200MB of RAM. Its lack of unnecessary background processes reduces disk I/O contention on mechanical hard drives, which is vital when multiple VMs are swapping.
  • Streamlined Architecture: Debian avoids using the snapd daemon and its loop-device overhead. This method keeps the host’s kernel routing table and mount points organised, simplifying the troubleshooting of VM network bridges.
  • Universal Binary Compatibility: Debian uses the standard glibc library. This compatibility allows enterprise virtualisation tools and guest agents to run natively without extra layers. As a result, Debian can use Ubuntu’s software ecosystem while minimising the associated service overhead.
Debian Cons
  • Manual Installation: Debian’s installer requires careful decisions about partitioning and mirror selection, as misconfigurations can cause issues.
  • Firmware Hurdles: Although hardware detection has improved, proprietary firmware for older network and WiFi cards may still need to be manually included from the “non-free” repository during initial setup (although from Debian 12 onwards this is less of a problem than it used to be).
  • Conservative Packaging: Debian prioritises stability over the latest software versions. Accessing the latest QEMU or libvirt features occasionally require the use of the backports repository.

Ubuntu Linux

Ubuntu is the most popular Linux distribution used in cloud and server environments. It is often the default choice referenced in DevOps and Kubernetes documentation. Commercially supported by Canonical, Ubuntu has a regular release schedule and offers long-term support versions suitable for both enterprise and home users.

Ubuntu Pros
  • Native Documentation Sync: Most KVM and libvirt guides are written with Ubuntu in mind. File paths and package names at the host level will closely align with the most common reference materials.
  • Extensive Driver Support: Ubuntu’s Hardware Enablement (HWE) stack includes proprietary drivers by default. This guarantees that niche hardware components on older PCs will work right away, eliminating the need for manual troubleshooting.
  • Commercial Maturity: Ubuntu provides a regular release schedule and strong community support for troubleshooting complex virtualisation scenarios such as GPU passthrough and nested virtualisation.
Ubuntu Cons
  • Service Overhead: The inclusion of snapd, cloud-init and automated update daemons result in an idle footprint of 300–400 MB. On hardware with low RAM, this consumes vital memory before a single VM starts.
  • Network Abstraction: Ubuntu uses the Netplan abstraction for network configuration. This adds a layer of YAML complexity as opposed to the direct control of the traditional /etc/network/interfaces file.
  • I/O Latency: Background services like unattended-upgrades can trigger sudden changes in disk activity. On older mechanical drives, this can cause guest VMs to experience latency spikes during high-load operations.

OS Decision

For my 2011 PC Kubernetes lab, Debian is the optimal choice. While Alpine is leaner and Ubuntu is more automated, Debian provides the best performance-to-stability ratio for a 2011-era PC acting as a VM host.

CPU and I/O Contention

Debian offers a more stable environment with fewer background processes compared to Ubuntu. This is particularly useful for a PC with an Intel i3-2120 dual-core processor, as it reduces the number of running daemons and minimises disk activity. This will improve VM performance and Kubernetes cluster stability.

Fastest Success Path

Debian occupies a sweet spot between Ubuntu’s excess and Alpine’s frugality, letting me set up a Kubernetes lab quickly without making significant changes to the OS.

Hypervisor Consistency

A hypervisor host functions best on a reliable foundation. Debian’s stable release offers fixed package versions and a cautious update policy, ensuring that the KVM toolchain, network bridge drivers and configurations stay consistent. This stability prevents version drift that could disrupt networking or storage for the guest VMs.

Testing Old PC

In this section, I conduct a smoke test of the 2011 PC and Debian to check that my repairs were successful and that the system is stable enough for virtualisation and Kubernetes.

Hard Drive Test

The first test is the simplest. When activated, the PC immediately detects the replacement hard drive. It then uses my Debian USB to run the installer, progressing through each stage without issue.

I’m installing Debian 13 – the most recent version, released on 9 August 2025. For those curious, this video shows a typical Debian 13 installation:

I mentioned earlier that the installer requires some careful decisions. One of these is the choice of Debian desktop environment:

2026 01 30 10 43 04 DebianSoftwareSelection

There are several choices here that control how minimal or feature-rich the desktop is. Many guides recommended the bare minimum SSH server to maximise the resources available for Kubernetes. I chose to install Xfce – a lightweight GUI designed for older or limited-resource hardware. This will let me see the various Debian backend files and folders in a familiar Windows-like environment, helping me understand Kubernetes’ inner workings.

This video explores several Debian desktop environments, including Xfce:

RAM Test

Next, let’s assess the RAM. This was partially verified during the hard drive test, as faulty RAM would have either caused issues during the Debian installation or prevented the PC from booting outright. But how much RAM is actually available?

With a successful Debian install, I can run free -h to find out! This shows 15GB memory with 14GB free:

2026 01 31 16 39 40 free

This is…interesting given the PC’s stated maximum of 8GB! Indeed, running dmidecode to query the BIOS states that the maximum capacity is 8GB:

2026 01 31 16 37 06 dmidecode

Having researched, there are two main reasons for this:

  • Outdated BIOS Info: The motherboard’s DMI table, which dmidecode uses for information, was created when 4GB sticks were the known maximum limit, and HP never updated it.
  • Ambiguous Definition: Sometimes “Max Capacity” refers to the per-slot maximum rather than the entire board‘s total. And there are two devices (slots), so 2x 8GB is 16GB.

In any case, free -h can be trusted so I do indeed have 16GB RAM!

I can also use lscpu to verify how many cores the PC has. The Intel Core i3-2120 is a dual-core processor with two physical cores. It features Intel Hyper-Threading, allowing each core to present two threads (logical processors) to the OS, resulting in a total of four threads.

2026 01 31 16 40 22 lscpu

Internet Test

Finally, let’s test the internet. Running ping -c 4 8.8.8.8 in the terminal returned packets successfully, and running sudo apt update fetched the latest package list:

2026 03 14 wifitest

With everything set up, there’s just one step remaining to convert my 2011 PC into a Kubernetes lab.

Configure Virtualisation

In this section, I outline the steps taken to install the software needed for the 2011 PC to function as a hypervisor in my Kubernetes lab. Please note that I do not claim this to be the definitive or best approach; it simply reflects my process for getting everything up and running on this occasion.

Installing Libraries

It’s now time to equip Debian with the necessary tools for virtualisation. This resulted in a fairly large install command, with each component contributing to my lab’s stability and manageability.

Bash
sudo apt install -y \
  qemu-kvm \
  bridge-utils \
  libvirt-daemon-system \
  libvirt-clients \
  virtinst \
  libosinfo-bin \
  virt-manager \
  virt-viewer

So, what is all this actually doing? Let’s examine in stages.

Core Components

qemu-kvm: This forms the core of the entire system. KVM is a Linux kernel module that lets a CPU run VMs at speeds nearly matching those of physical hardware. QEMU manages the rest of the hardware emulation, providing the virtual motherboard, disk controllers and network cards that the guest OS recognises as genuine. Together, these components turn a PC into a platform capable of hosting multiple VMs.

bridge-utils: While VMs are isolated by default, bridge-utils creates a virtual switch that links them directly to a physical network. This is crucial for a Kubernetes lab, as it allows each node to obtain its own IP address and route traffic through the Raspberry Pi bridge. By using the host’s Ethernet connection in this way, the VMs gain the stable internet access required for cluster orchestration and updates.

Lifecycle & Control

libvirt-daemon-system: This package operates as a daemon responsible for managing VM lifecycles. It monitors active nodes, controls startup and shutdown processes, and ensures that a VM crash does not impact the entire system.

libvirt-clients: This provides the command-line tools needed to interact with libvirt-daemon-system, most notably including virsh – an industry-standard utility for managing VMs via the terminal. It can check statuses, modify configurations and reboot nodes from the terminal.

Provisioning & Optimisation

virtinst: This package yields virt-install, a command-line tool for creating VMs. Instead of going through a wizard multiple times, one command now specifies the CPU, RAM, and disk for multiple nodes. This makes setting up Kubernetes nodes quick and repeatable.

libosinfo-bin: Virtualisation works best when the host has detailed knowledge of the guest OS. This tool serves as a database that guides the hypervisor in using appropriate high-performance drivers for each VM. This technique allows the main CPU to conserve its cycles by reducing reliance on inefficient hardware emulation.

Monitoring & Troubleshooting

virt-manager: Although the primary goal is to manage the VMs through the terminal, I also want a GUI as I familiarise myself with the system. Virt-manager offers a graphical interface to monitor CPU and RAM usage in real time.

virt-viewer: If a VM fails to boot or loses its network connection, SSH access is impossible. This package opens a window that displays the VM’s console, allowing users to troubleshoot the boot process or fix a broken config file as if accessing it directly.

Testing Virtualisation

Having installed everything, I now need to test that the hypervisor works. On Debian, installing the packages doesn’t always mean the services are active and ready to run. So firstly I need to start the virtualisation daemon manually:

Bash
sudo systemctl enable --now libvirtd

And then check virtualisation is running:

Bash
sudo systemctl status libvirtd --no-pager

It should show a green active (running) status. But…uh oh.

2026 02 05 14 51 24 MissingVirtualisation

This error means virtualisation isn’t enabled. And that means a trip to the PC’s BIOS settings!

2026 02 07 18 32 36 BIOSVirt

With VTx enabled and a reboot later, the libvirtd service is now active:

2026 02 05 14 54 40 VirtOK

Running kvm-ok (installed via cpu-checker)also shows the PC can now host hardware-accelerated KVM virtual machines. I can now run virt-manager to load the GUI:

2026 02 05 14 55 52 kvm

This loads the Virtual Machine Manager, from which I can create and interact with my VMs via a GUI:

2026 02 05 14 56 29 VirtualMachineManager

My 2011 PC can now spin up some Kubernetes nodes!

In terms of next steps, I’ll use the Kubernetes documentation to turn the 2011 PC into a multi-node cluster. The CKA is open-book, and getting familiar with the documentation will be a massive help during the exam!

Summary

In this post, I documented the journey of resurrecting a dusty old 2011 PC into a functional Kubernetes lab to prepare for my CKA exam.

This experience was a nice blend of familiar and new. While I’ve worked extensively with Linux, I’ve not gone into the weeds with various distros before. Nor have I built my own NAT router! Ultimately, this was a fun and engaging project and I’m looking forward to kicking the tyres of my new lab!

Like this post? Click the button below for links to contact, socials, projects and sessions:

SharkLinkButton 1

Thanks for reading ~~^~~

Categories
Data & Analytics

Python Data Validation And Observability As Code With Pydantic

In this post, I use the Pydantic Python library to create data validation and observability processes for my Project Wolfie iTunes data.

Introduction

Data validation is a crucial component of any data project. It ensures that data is accurate, consistent and reliable. It verifies that data meets set criteria and rules to maintain its quality, and stops erroneous or unreliable information from entering downstream systems. I’ve written about it, scripted it and talked about it.

Validation will be a crucial aspect of Project Wolfie. It is an ongoing process that should occur from data ingestion to exposure, and should be automated wherever possible. Thankfully, most data processes within Project Wolfie are (and will be) built using Python, which provides several libraries to simplify data validation. These include Pandera, Great Expectations and the focus of this post – Pydantic (specifically, version 2).

Firstly, I’ll explore the purpose and benefits of Pydantic. Next, I’ll import some iTunes data and use it to explore key Pydantic validation concepts. Finally, I’ll explore how Pydantic handles observability and test its findings. The complete code will be in a GitHub repo.

Let’s begin!

Introducing Pydantic

This section introduces Pydantic and examines some of its benefits.

About Pydantic

Pydantic is an open-source data validation Python library. It uses established Python notation and constructs to define data structures, types and constraints. These can then validate the provided data, generating clear error messages when issues occur.

Pydantic is a widely used tool for managing application settings, validating API requests and responses, and streamlining data transfer between Python objects and formats like JSON. By integrating both existing and custom elements, it offers a powerful and Pythonic method for ensuring data quality and consistency within projects. This makes data handling in Python more reliable and reduces the likelihood of errors through its intuitive definition and validation processes.

Pydantic Benefits

Pydantic’s benefits are thoroughly documented, and the ones I want to highlight here are:

Intuitive: Pydantic’s use of type hints, functions and classes fits well with my current Python skill level, so I can focus on learning Pydantic without also having to explore unfamiliar Python concepts.

Fast: Pydantic’s core validation logic is written in Rust, which enables rapid development, testing, and validation. This speed has contributed towards…

Well-Supported: Pydantic has extensive community use and support from organisations like Anthropic, Netflix and OpenAI, as well as popular Python libraries like Airflow, FastAPI and LangChain. It also has extensive AWS Lambda support via user-configurable artefacts and the community-managed Powertools for AWS Lambda (Python)‘s Parser utility.

Preparation

Before I can start using Pydantic, I need some data. This section examines the data I am using and how I prepare it for Pydantic.

iTunes Data

Firstly, let’s extract some data from iTunes. I create iTunes Export files using the iTunes > Export Playlist command. Apple has documented this, but WikiHow’s documentation is more illustrative. The export file type choices are…interesting. The one closest to matching my needs is the txt format, although the files are technically tab-separated files (TSVs).

iTunes Exports contain many metadata columns. I’m not including them all here (after all, this is a Pydantic post not an iTunes one), but I will be using the following subset (using my existing metadata definitions):

Metadata TypeColumn NameData TypePurpose
TechnicalAlbumStringTrack key as Camelot Notation*
TechnicalLocationStringTrack file path
TechnicalTrack NumberIntegerTrack BPM*
DescriptiveArtistStringTrack artist(s)
DescriptiveGenreStringTrack genre
DescriptiveNameStringTrack name and mix
DescriptiveWorkStringPublishing record label
DescriptiveYearIntegerTrack release year
InteractionMy RatingIntegerTrack personal rating

Note that the starred Album and Track Number columns have purposes that differ from the column names. The reasons for this are…not ideal.

  • Track Number contains BPM data as, although iTunes does have a BPM column, it isn’t included in the exports. And the exports can’t be customised! To include BPMs in an export, I had to repurpose an existing column.

Great. But that’s not as bad as…

  • Album contains musical keys, as iTunes doesn’t even have a key column, despite MP3s having a native Initial Key metadata field! Approaches to dealing with this vary – I chose to use another donor column. I’ll explain Camelot Notations later on.

That’s enough about the iTunes data for now – I’ll go into more detail in future Project Wolfie posts. Now let’s focus on getting this data into memory for Python.

Data Capture

Next, let’s get the iTunes data into memory. Starting with a familiar library…

pandas

I’ll be using pandas to ingest the iTunes data. This is a well-established and widely supported module. It also has its own data validation functions and will assist with issues like handling spaces in column names.

While iTunes files aren’t CSVs, the pandas read_csv function can still read their data into a DataFrame. It needs some help though – the delimiter parameter must be \t to identify the tabs’ delimiting status.

So let’s read the iTunes metadata into memory and…

Python
df = pd.read_csv(csv_path, delimiter='\t')

>> UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte

Oh. pandas can’t read the file. The error says it’s trying the utf-8 codec, so the export must be using something else. Fortunately, there’s another Python library that can help!

charset_normalizer

charset_normalizer is an open-source encoding detector. It determines the encoding of a file or text and records the result. It’s related to the older chardet library but is faster, has a more permissive MIT license and supports more encodings.

Here, I’m using charset_normalizer.detect in a detect_file_encoding function to detect the export’s codec:

Python
def detect_file_encoding(file_path: Path) -> str:
    with open(file_path, 'rb') as file:
        raw_data = file.read()
        
    detection_result = charset_normalizer.detect(raw_data)
    return detection_result['encoding'] or 'utf-8'

In which:

  • I define a detect_file_encoding function that expects a filepath and returns a string.
  • detect_file_encoding opens the file, reads the data and stores it as raw_data.
  • charset_normalizer detects raw_data‘s codec and stores this as detection_result.
  • detect_file_encoding returns either the successfully detected codec, or the common utf-8 codec if the attempt fails.

I can then pass the export’s filepath to the detect_file_encoding function, capture the results as encoding and pass this as a parameter to pandas.read_csv:

Python
encoding = detect_file_encoding(csv_path)
    
df = pd.read_csv(csv_path, encoding=encoding, delimiter='\t')

>> Loaded 4407 rows

There’s one more action to take before moving on. Some columns contain spaces. This will become a problem as spaces are not allowed in Python identifiers!

As the data is now in a pandas DataFrame, I can use pandas.DataFrame.rename to remove these spaces:

Python
df = df.rename(columns={
        'Track Number': 'TrackNumber',
        'My Rating': 'MyRating'
    })

The metadata is now ready for Pydantic.

Installing Pydantic

Finally, let’s install Pydantic. This process is fully documented. My preferred method is via pip install in a local virtual environment:

Python
pip install pydantic

And then importing Pydantic into my script:

Python
import pydantic

Now I can start using Pydantic.

Pydantic Data Models

In this section, I tell Pydantic about my data model and the types of data it should expect for validation.

Introducing BaseModel

At the core of Pydantic is the BaseModel class – used for defining data models. Every Pydantic model inherits from it, and by doing so gains features like type enforcement, automatic data parsing and built-in validation.

By subclassing BaseModel, a schema for the data is defined using standard Python type hints. Pydantic uses these hints to validate and convert input data automatically.

Let’s explore BaseModel by creating a new Track class.

Creating A Track Class

Pydantic supports standard library types like string and integer. This reduces Pydantic’s learning curve and simplifies integration into existing Python processes.

Here are the very beginnings of my Track data model. I have a new Track class inheriting from Pydantic’s BaseModel, and a Name field with string data type:

Python
class Track(BaseModel):
    Name: str

Next, I add a Year field with integer data type:

Python
class Track(BaseModel):
    Name: str
    Year: int

And so on for each field I want to validate with Pydantic:

Python
class Track(BaseModel):
    Name: str
    Artist: str
    Album: str
    Work: str
    Genre: str
    TrackNumber: int
    Year: int
    MyRating: int
    Location: str

Now, if any field is missing or has the wrong type, Pydantic will raise a ValidationError. But there’s far more to Pydantic data types than this…

Defining Special Data Types

Where no standards exist or where validation rules are more complex to determine, Pydantic offers further type coverage. These include:

One of my Track fields will immediately benefit from this:

Python
class Track(BaseModel):
    Location: str

Currently, my Location field validation is highly permissive. It will accept any string. I can improve this using Pydantic’s FilePath data type:

Python
class Track(BaseModel):
    Location: FilePath

Now, Pydantic will check that the given location is a path that exists and links to a valid file. No custom code; no for loops – the FilePath type handles everything for me.

So I now have data type validation in my Pydantic data model. What else can I have?

Pydantic Built-In Validation

This section explores the native data validation features of Pydantic, including field annotation and constraints.

Introducing Field

In Pydantic models, data attributes are typically defined using Python type hints. The Field function enables further customisation like constraints, schema metadata and default values.

While type hints define what kind of data is allowed, Field defines how that data should behave, what happens if it’s missing and how it should be documented. It adds clarity to models and helps Pydantic enforce stricter rules.

Let’s run through some examples.

Custom Schema Metadata

One of the challenges in creating data pipelines is that the data fields can sometimes be unclear or difficult to explain. This can cause confusion and delay when building ETLs, examining repos and interacting with code.

Field helps here by adding custom fields to annotate data within Pydantic classes. Examples include description:

Python
class Track(BaseModel):
    Name: str = Field(
        description="Track's name and mix.")

And examples:

Python
class Track(BaseModel):
    Name: str = Field(
        description="Track's name and mix.",
        examples=["Track Title (Original Mix)", "Track Title (Extended Mix)"])

Using these throughout my Track class simplifies the code and reduces context switching:

Python
class Track(BaseModel):
    Name: str = Field(
        description="Track's name and mix.",
        examples=["Track Title (Original Mix)", "Track Title (Extended Mix)"])
    
    Artist: str = Field(
        description="The artist(s) of the track.",
        examples=["Above & Beyond", "Armin van Buuren"])
    
    Album: str = Field(
        description="Track's Camelot Notation indicating the key.",
        examples=["01A-Abm", "02B-GbM"])
    
    Work: str = Field(
        description="The record label that published the track.",
        examples=["Armada Music", "Anjunabeats"])
    
    Genre: str = Field(
        description="Track's musical genre.",
        examples=["Trance", "Progressive House"])
    
    TrackNumber: int = Field(
        description="Track's BPM (Beats Per Minute).",
        examples=[130, 140])
    
    Year: int = Field(
        description="Track's release year.",
        examples=[1998, 2004])
    
    MyRating: int = Field(
        description="Personal Rating.  Stars expressed as 0, 20, 40, 60, 80, or 100",
        examples=[60, 80])
    
    Location: FilePath = Field(
        description="Track's Location on the filesystem.",
        examples=[r"C:\Users\User\Music\iTunes\TranquilityBase-GettingAway-OriginalMix.mp3"])

This is especially useful for Album and TrackNumber given their unique properties.

Field Constraints

Field can also constrain the data that a class accepts. This includes string constraints:

  • max_length: Maximum length of the string.
  • min_length: Minimum length of the string.
  • pattern: A regular expression that the string must match.

and numeric constraints:

  • ge & le – greater than or equal to/less than or equal to
  • gt & lt – greater/less than
  • multiple_of – multiple of a given number

Constraints can also be combined as needed. For example, iTunes exports record MyRating values in increments of 20, where 1 star is 20 and 2 stars are 40, rising to the maximum 5 stars being 100.

I can express this within the Track class as:

Python
class Track(BaseModel):
    MyRating: int = Field(
        description="Personal Rating.  Stars expressed as 0, 20, 40, 60, 80, or 100",
        examples=[60, 80],
        ge=20,
        le=100,
        multiple_of=20)

Here, MyRating must be greater than or equal to 20 (ge=20), less than or equal to 100 (le=100), and must be a multiple of 20 (multiple_of=20).

I can also parameterise these constraints using variables instead of hard-coded values:

Python
ITUNES_RATING_RAW_LOWEST = 20
ITUNES_RATING_RAW_HIGHEST = 100

class Track(BaseModel):
    MyRating: int = Field(
        description="Personal Rating.  Stars expressed as 0, 20, 40, 60, 80, or 100",
        examples=[60, 80],
        ge=ITUNES_RATING_RAW_LOWEST,
        le=ITUNES_RATING_RAW_HIGHEST,
        multiple_of=20)

This property lets me use Pydantic with other Python libraries. Here, my Year validation checks for years greater than or equal to 1970 and less than or equal to the current year (using the datetime library):

Python
YEAR_EARLIEST = 1970
YEAR_CURRENT = datetime.datetime.now().year

class Track(BaseModel):
    Year: int = Field(
        description="Track's release year.",
        examples=[1998, 2004],
        ge=YEAR_EARLIEST,
        le=YEAR_CURRENT)

No track in the collection should exist beyond the current year – this constraint will now update itself as time passes.

Having applied other constraints, my Track class looks like this:

Python
class Track(BaseModel):
    """Pydantic model for validating iTunes track metadata."""
    
    Name: str = Field(
        description="Track's name and mix type.",
        examples=["Track Title (Original Mix)", "Track Title (Extended Mix)"])
    
    Artist: str = Field(
        description="The artist(s) of the track.",
        examples=["Above & Beyond", "Armin van Buuren"])
    
    Album: str = Field(
        description="Track's Camelot Notation indicating the key.",
        examples=["01A-Abm", "02B-GbM"])
    
    Work: str = Field(
        description="The record label that published the track.",
        examples=["Armada Music", "Anjunabeats"])
    
    Genre: str = Field(
        description="Track's musical genre.",
        examples=["Trance", "Progressive House"])
    
    TrackNumber: int = Field(
        description="Track's BPM (Beats Per Minute).",
        examples=[130, 140],
        ge=BPM_LOWEST,
        le=BPM_HIGHEST)
    
    Year: int = Field(
        description="Track's release year.",
        examples=[1998, 2004],
        ge=YEAR_EARLIEST,
        le=YEAR_CURRENT)
    
    MyRating: int = Field(
        description="Personal Rating. Stars expressed as 0, 20, 40, 60, 80, or 100",
        examples=[60, 80],
        ge=ITUNES_RATING_RAW_LOWEST,
        le=ITUNES_RATING_RAW_HIGHEST,
        multiple_of=20)
    
    Location: FilePath = Field(
        description="Track's Location on the filesystem.",
        examples=[r"C:\Users\User\Music\iTunes\AboveAndBeyond-AloneTonight-OriginalMix.mp3"])

This is already very helpful. Next, let’s examine my custom requirements.

Pydantic Custom Validation

This section discusses how to create custom data validation using Pydantic. I will outline what the requirements are, and then examine how these validations are defined and implemented.

Introducing Decorators

In Python, decorators modify or enhance the behaviour of functions or methods without changing their actual code. Decorators are usually written using the @ symbol followed by the decorator name, just above the function definition:

Python
@my_decorator
def my_function():
    ...

For example, consider this logger_decorator function:

Python
def logger_decorator(func):
    def wrapper():
        print(f"Running {func.__name__}...")
        func()  # Execute the supplied function
        print("Done!")
    return wrapper

This function takes another function (func) as an argument, printing a message before and after execution. If the logger_decorator function is then used as a decorator when running this greet function:

Python
@logger_decorator
def greet():
    print("Hello, world!")

greet()

Python will add the logging behaviour of logger_decorator without modifying greet:

Python
Running greet...
Hello, world!
Done!

Introducing Field Validators

In addition to the built-in data validation capabilities of Pydantic, custom validators with more specific rules can be defined for individual fields using Field Validators. These use the field_validator() decorator, and are declared as class methods within a class inheriting from Pydantic’s BaseModel.

Here’s a basic example using my Track model:

Python
class Track(BaseModel):
    Name: str = Field(
        description="Track's name and mix.",
        examples=["Track Title (Original Mix)", "Track Title (Extended Mix)"]
    )

    @field_validator("Name")
    @classmethod
    def validate_name(cls, value):
        # custom validation logic here
        return value

Where:

  • @field_validator("Name") tells Pydantic to use the function to validate the Name field.
  • @classmethod lets the validator access the Track class (cls).
  • The validator executes the validate_name function with the field value (in this case Name) as input, performs the checks and must either:
    • return the validated value, or
    • raise a ValueError or TypeError if validation fails.

Let’s see this in action.

Null Checks

Firstly, let’s perform a common data validation check by identifying empty fields. I have two variants of this – one for strings and another for numbers.

The first – validate_non_empty_string – uses pandas.isna to catch missing values and strip() to catch empty strings. This field validator applies to the Artist, Work and Genre columns:

Python
    @field_validator("Artist", "Work", "Genre")
    @classmethod
    def validate_non_empty_string(cls, value, info):
        """Validate that a string field is not empty."""
        if pd.isna(value) or str(value).strip() == "":
            raise ValueError(f"{info.field_name} must not be null or empty")
        return value

The second – validate_non_null_numeric – checks the TrackNumber, Year and MyRating numeric columns for empty values using pandas.isna:

Python
    @field_validator("TrackNumber", "Year", "MyRating", mode="before")
    @classmethod
    def validate_non_null_numeric(cls, value, info):
        """Validate that a numeric field is not null."""
        if pd.isna(value):
            raise ValueError(f"{info.field_name} must not be null")
        return value

Also, it uses Pydantic’s before validator (mode="before"), ensuring the data validation happens before Pydantic coerces types. This catches edge cases like "" or "NaN" before they become None or float("nan") values.

Character Check

Now let’s create a validator for something a little more challenging to define. All tracks in my collection follow a Track Name (Mix) schema. This can take many forms:

  • Original track: Getting Away (Original Mix)
  • Remixed track: Shapes (Oliver Smith Remix)
  • Updated remixed track: Distant Planet (Menno de Jong Interpretation) (2020 Remaster)
  • …and many more variants.

But generally, there should be at least one instance of text enclosed by parentheses. However, some tracks have no remixer and are released with just a title:

  • Getting Away
  • Shapes
  • Distant Planet

This not only looks untidy (eww!), but also breaks some of my downstream automation that expects the Track Name (Mix) schema. So any track without a remixer gets (Original Mix) added to the Name field upon download:

  • Getting Away (Original Mix)
  • Shapes (Original Mix)
  • Distant Planet (Original Mix)

Expressing this is possible with RegEx, but I can make a more straightforward and more understandable check with a field validator:

Python
    @field_validator("Name")
    @classmethod
    def validate_name(cls, value):
        if pd.isna(value) or str(value).strip() == "":
            raise ValueError("Name must not be null or empty")
        
        value_str = str(value)
        if '(' not in value_str:
            raise ValueError("Name must contain an opening parenthesis '('")
        if ')' not in value_str:
            raise ValueError("Name must contain a closing parenthesis ')'")
        return value

This validator checks that the value isn’t empty and then performs additional checks for parentheses. This could be one check, but having it as two checks improves log readability (insert foreshadowing – Ed). I could also have added Name to the validate_non_empty_string validation, but this way I have all my Name checks in the same place.

Parameterised Checks

Like constraints, field validators can also be parameterised. Let’s examine Album.

As iTunes exports can’t be customised, I use Album for a track’s Camelot Notation. These are based on the Camelot WheelMixedInKey‘s representation of the Circle Of Fifths. DJs generally favour Camelot Notation as it is simpler than traditional music notation for human understanding and application sorting.

Importantly, there are only twenty-four possible notations:

For example:

  • 1A (A-Flat Minor)
  • 6A (G Minor)
  • 6B (B-Flat Major)
  • 10A (B Minor)

So let’s capture these values in a CAMELOT_NOTATIONS list:

Python
CAMELOT_NOTATIONS = {
    '01A-Abm', '01B-BM', '02A-Ebm', '02B-GbM', '03A-Bbm', '03B-DbM',
    '04A-Fm', '04B-AbM', '05A-Cm', '05B-EbM', '06A-Gm', '06B-BbM',
    '07A-Dm', '07B-FM', '08A-Am', '08B-CM', '09A-Em', '09B-GM',
    '10A-Bm', '10B-DM', '11A-Gbm', '11B-AM', '12A-Dbm', '12B-EM'
}

(Note the leading zeros. Without them, iTunes sorts the Album column as (10, 11, 12, 1, 2, 3…) – you can imagine how I felt about that – Ed)

Next, I pass the CAMELOT_NOTATIONS list to an Album field validator that checks if the given value is in the list:

Python
    @field_validator("Album")
    @classmethod
    def validate_album(cls, value):
        if pd.isna(value) or str(value).strip() == "":
            raise ValueError("Album must not be null or empty")
        
        if str(value) not in CAMELOT_NOTATIONS:
            raise ValueError(f"Album must be a valid Camelot notation: {value} is not in the valid list")
        return value

Pydantic now fails any value not found in the CAMELOT_NOTATIONS list.

Now I have my validation needs fully covered. What observability does Pydantic give me over these data validation checks?

Pydantic Observability

In this section, I assess and adjust the default Pydantic observability abilities to ensure my data validation is accurately recorded.

Default Output

Pydantic automatically generates data validation error messages if validation fails. These detailed messages provide a structured overview of the issues encountered, including:

  • The index of the failing input (e.g., a DataFrame row number).
  • The model class where the error occurred.
  • The field name that failed validation.
  • A human-readable explanation of the issue.
  • The offending input value and its type.
  • A direct link to relevant documentation for further guidance.

Here’s an example of Pydantic’s output when a string field receives a NaN value:

Python
Row 2353: 1 validation error for Track
Work
  Input should be a valid string [type=string_type, input_value=nan, input_type=float]
    For further information visit https://errors.pydantic.dev/2.11/v/string_type

In this example:

  • Row 2353 indicates the problematic input row.
  • Track is the Pydantic model where validation failed.
  • Work is the failing field.
  • Pydantic detects that the input is nan (a float) and not a valid string.
  • Pydantic provides a URL to the string_type documentation.

Here’s another example, this time for a MyRating error:

Python
Row 3040: 1 validation error for Track
MyRating
  Value error, MyRating must not be null [type=value_error, input_value=nan, input_type=float]
    For further information visit https://errors.pydantic.dev/2.11/v/value_error

In this case, a field validator raised a ValueError because MyRating must not be null.

Pydantic’s error reporting is clear and actionable, making it suitable for debugging and systemic data validation tasks. However, for larger datasets or more user-friendly outputs (such as reports or UI feedback), further customisation is helpful, such as…

Terminal Output Customisation

As good as Pydantic’s default output is, it’s not that human-readable. For example, in this Terminal output I have no idea which tracks are on rows 2353, 2495 and 3040:

Plaintext
Row 2353: 1 validation error for Track
Work
  Input should be a valid string [type=string_type, input_value=nan, input_type=float]
    For further information visit https://errors.pydantic.dev/2.11/v/string_type
    
Row 2495: 1 validation error for Track
Work
  Input should be a valid string [type=string_type, input_value=nan, input_type=float]
    For further information visit https://errors.pydantic.dev/2.11/v/string_type
    
Row 3040: 1 validation error for Track
MyRating
  Value error, MyRating must not be null [type=value_error, input_value=nan, input_type=float]
    For further information visit https://errors.pydantic.dev/2.11/v/value_error

While I can find this out, it would be better to know at a glance. Fortunately, I can improve this when capturing the errors by appending the artist and name to each row of the errors object:

Python
except (ValidationError, ValueError) as e:
            artist = row['Artist'] if not pd.isna(row['Artist']) else "Unknown Artist"
            name = row['Name'] if not pd.isna(row['Name']) else "Unknown Name"
            errors.append((index, artist, name, str(e)))

Now, Artist and Name are added to each row:

Plaintext
Row 2353: Ben Stone - Mercure (Extended Mix): 1 validation error for Track
Work
  Input should be a valid string [type=string_type, input_value=nan, input_type=float]
    For further information visit https://errors.pydantic.dev/2.11/v/string_type
    
Row 2495: DJ Hell - My Definition Of House Music (Resistance D Remix): 1 validation error for Track
Work
  Input should be a valid string [type=string_type, input_value=nan, input_type=float]
    For further information visit https://errors.pydantic.dev/2.11/v/string_type
    
Row 3040: York - Reachers Of Civilisation (In Search Of Sunrise Mix): 1 validation error for Track
MyRating
  Value error, MyRating must not be null [type=value_error, input_value=nan, input_type=float]
    For further information visit https://errors.pydantic.dev/2.11/v/value_error

This makes it far easier to find the problematic files in my collection. As long as there aren’t many findings…

Creating An Error File

There are three main problems with Pydantic printing all data validation errors in the Terminal:

  • They don’t persist outside of the Terminal session.
  • The Terminal isn’t that easy to read when it’s full of text.
  • The Terminal may run out of space if there are a large number of errors.

So let’s capture the errors in a file instead. This write_error_report function generates a text-based error report from validation failures, saving it in a logs subfolder adjacent to the input file:

Python
def write_error_report(
    csv_path: Path, 
    field_error_details: Dict[str, List[str]],  
    sorted_fields: List[str]
) -> Path:

    timestamp = datetime.now().strftime("%Y%m%d-%H%M%S")
    logs_dir = csv_path.parent / "logs"
    logs_dir.mkdir(exist_ok=True)
    
    error_output_path = logs_dir / f"{timestamp}-PydanticErrors-{csv_path.stem}.txt"
    
    with open(error_output_path, 'w', encoding='utf-8') as f:
        f.write(f"Validation Error Report - {timestamp}\n")
        f.write("=" * 80 + "\n")

        for field in sorted_fields:
            messages = field_error_details.get(field, [])
            if messages:
                f.write(f"\n{field} Errors ({len(messages)}):\n")
                f.write("-" * 80 + "\n")
                for message in messages:
                    f.write(message + "\n\n")
    
    return error_output_path

Firstly, it constructs a timestamped filename using the original file’s stem (e.g., 20250529-142304-PydanticErrors-data.txt) and the logs subfolder, creating the latter if it doesn’t exist:

Python
    timestamp = datetime.now().strftime("%Y%m%d-%H%M%S")
    logs_dir = csv_path.parent / "logs"
    logs_dir.mkdir(exist_ok=True)
    
    error_output_path = logs_dir / f"{timestamp}-PydanticErrors-{csv_path.stem}.txt"

Next, Python orders the errors by the sorted_fields input, displays error counts per field and formats each error message with clear section dividers. A structured report listing all validation errors by field is saved in the logs subfolder:

Python
    with open(error_output_path, 'w', encoding='utf-8') as f:
        f.write(f"Validation Error Report - {timestamp}\n")
        f.write("=" * 80 + "\n")

        for field in sorted_fields:
            messages = field_error_details.get(field, [])
            if messages:
                f.write(f"\n{field} Errors ({len(messages)}):\n")
                f.write("-" * 80 + "\n")
                for message in messages:
                    f.write(message + "\n\n")

Finally, the filesystem path of the generated report is returned:

Python
    return error_output_path

When executed, the Terminal tells me the error file path:

Plaintext
Detailed error log written to: 20250513-133743-PydanticErrors-iTunes-Elec-Dance-Club-Main.txt

And stores the findings in a local txt file, grouped by error type for simpler readability:

Plaintext
Validation Error Report - 20250513-133743
================================================================================

MyRating Errors (5):
--------------------------------------------------------------------------------
Row 3040: York - Reachers Of Civilisation (In Search Of Sunrise Mix): 1 validation error for Track
MyRating
  Value error, MyRating must not be null [type=value_error, input_value=nan, input_type=float]
    For further information visit https://errors.pydantic.dev/2.11/v/value_error

Work Errors (22):
--------------------------------------------------------------------------------
Row 223: Dave Angel - Artech (Original Mix): 1 validation error for Track
Work
  Input should be a valid string [type=string_type, input_value=nan, input_type=float]
    For further information visit https://errors.pydantic.dev/2.11/v/string_type

Adding A Terminal Summary

Finally, I created a Terminal summary of Pydantic’s findings:

Python
print("\nValidation Summary:\n")
sorted_fields = sorted(Track.model_fields.keys())

for field in sorted_fields:
  count = error_analysis['counts'].get(field, 0)
  print(f"{field} findings: {count}")

This shows feedback after each execution:

Plaintext
Validation Summary:

Album findings: 0
Artist findings: 0
Genre findings: 0
Location findings: 0
MyRating findings: 5
Name findings: 1
TrackNumber findings: 0
Work findings: 22
Year findings: 0

Now, let’s ensure everything works properly!

Testing Pydantic

In this section, I test that my Pydantic data validation and observability processes are working correctly using iTunes export files and pytest unit tests.

Recent File Test

The first test used a recent export from the end of April 2025. Here is the Terminal output:

Plaintext
Processing file: iTunes-Elec-Dance-Club-Main-2025-04-28.txt
Reading iTunes-Elec-Dance-Club-Main-2025-04-28.txt with detected encoding UTF-16
Loaded 4407 rows
Validated 4379 rows
Found 28 errors!

Validation Summary for iTunes-Elec-Dance-Club-Main-2025-04-28.txt:
Album errors: 0
Artist errors: 0
Genre errors: 0
Location errors: 0
MyRating errors: 5
Name errors: 1
TrackNumber errors: 0
Work errors: 22
Year errors: 0

Detailed error log written to: 20250521-164324-PydanticErrors-iTunes-Elec-Dance-Club-Main-2025-04-28.txt

Good first impressions – the 4407 row count matches the export file, the summary is shown in the Terminal and an error log is created. So what’s in the log?

Firstly, five tracks have no MyRating values. For example:

Plaintext
MyRating Errors (5):
--------------------------------------------------------------------------------
Row 558: Reel People Feat Angela Johnson - Can't Stop (Michael Gray Instrumental Remix): 1 validation error for Track
MyRating
  Value error, MyRating must not be null [type=value_error, input_value=nan, input_type=float]
    For further information visit https://errors.pydantic.dev/2.11/v/value_error

This is correct, as this export was created when I added some new tracks to my collection.

Next, one track has a Name issue:

Plaintext
Name Errors (1):
--------------------------------------------------------------------------------
Row 1292: The Prodigy - Firestarter (Original Mix}: 1 validation error for Track
Name
  Value error, Name must contain a closing parenthesis ')' [type=value_error, input_value='Firestarter (Original Mix}', input_type=str]
    For further information visit https://errors.pydantic.dev/2.11/v/value_error

This one confused me at first, until I looked at the error more closely and realised the closing parenthesis is wrong! } is used instead of )! This is why my validate_name field validator has separate checks for each character – it makes it easier to understand the results!

Finally, twenty-two tracks are missing record label metadata in Work:

Plaintext
Work Errors (22):
--------------------------------------------------------------------------------
Row 223: Dave Angel - Artech (Original Mix): 1 validation error for Track
Work
  Input should be a valid string [type=string_type, input_value=nan, input_type=float]
    For further information visit https://errors.pydantic.dev/2.11/v/string_type

This means some tracks are missing full metadata. This won’t break any downstream processes as I have no reliance on this field. That said, it’s good to know about this in case my future needs change.

Older File Test

The next test uses an older file from March 2025. Let’s see what the Terminal says this time…

Plaintext
Processing file: iTunes-AllTunesMaster-2025-03-01.txt      
Reading iTunes-AllTunesMaster-2025-03-01.txt with detected encoding UTF-16
Loaded 4381 rows
Validated 0 rows
Found 4381 errors!

Validation Summary for iTunes-AllTunesMaster-2025-03-01.txt:

Album errors: 0
Artist errors: 0
Genre errors: 0
Location errors: 4381
MyRating errors: 0
Name errors: 1
TrackNumber errors: 2
Work errors: 17
Year errors: 0
Detailed error log written to: 20250521-164322-PydanticErrors-iTunes-AllTunesMaster-2025-03-01.txt

There are fewer rows here – 4381 vs 4407. This is correct, as my collection was smaller in March. But no rows were validated successfully!

I don’t have to go far to find out why:

Plaintext
Location Errors (4381):
--------------------------------------------------------------------------------
Row 0: Ariel - A9 (Original Mix): 1 validation error for Track
Location
  Path does not point to a file [type=path_not_file, input_value='C:\\Users\\User\\Folder...riel-A9-OriginalMix.mp3', input_type=str]

All the location checks failed. But this is actually a successful test!

In the time between these two exports, I reorganised my music collection. As a result, the file paths in this export no longer exist. Remember – the Location field uses the FilePath data type, which checks that the given paths exist and link to valid files. And these don’t!

The Name results are the same as the first test. This has been around for a while apparently…

Plaintext
Name Errors (1):
--------------------------------------------------------------------------------
Row 1292: The Prodigy - Firestarter (Original Mix}: 1 validation error for Track
Name
  Value error, Name must contain a closing parenthesis ')' [type=value_error, input_value='Firestarter (Original Mix}', input_type=str]
    For further information visit https://errors.pydantic.dev/2.11/v/value_error

There are also TrackNumber errors in this export:

Plaintext
TrackNumber Errors (2):
--------------------------------------------------------------------------------
Row 485: Andrew Bayer Feat Alison May - Brick (Original Mix): 2 validation errors for Track
TrackNumber
  Input should be greater than or equal to 100 [type=greater_than_equal, input_value=90, input_type=int]
    For further information visit https://errors.pydantic.dev/2.11/v/greater_than_equal

Two tracks have BPM values lower than the set range. Both files were moved during my reorganisation, but were included in this export at the time and therefore fail this validation check.

Finally, the Work errors are the same as the first test (although more have crept in since!):

Plaintext
Work Errors (17):
--------------------------------------------------------------------------------
Row 223: Dave Angel - Artech (Original Mix): 1 validation error for Track
Work
  Input should be a valid string [type=string_type, input_value=nan, input_type=float]
    For further information visit https://errors.pydantic.dev/2.11/v/string_type

Ultimately, both tests match expectations!

Unit Tests With Amazon Q

Finally, I wanted to include some unit tests for this project. Unit testing is always a good idea, especially in this context where I can verify function outputs and error generation without needing to create numerous test files.

I figured this was a good opportunity to test Amazon Q Developer and see what it came up with. I gave it a fairly basic prompt, using the @workspace context to allow Q access to my project’s entire workspace as context for its responses:

Plaintext
@workspace write unit tests for this script using pytest

I tend to use pytest for my Python testing, as I find it simpler and more flexible than Python’s standard unittest library.

Q promptly provided several reasonable tests in response. This initiated a half-hour exchange between us focused on calibrating the existing tests and creating new ones. To be fair to Q, my initial prompt was quite basic and could have been much more detailed.

Amongst Q’s tests was this one testing an empty Artist field:

Python
    @patch('pathlib.Path.exists')
    def test_empty_artist(self, mock_exists):
        """Test that an empty artist fails validation."""
        # Mock file existence check
        mock_exists.return_value = True
        
        invalid_track_data = {
            "Name": "Test Track (Original Mix)",
            "Artist": "",  # Empty artist
            "Album": "01A-Abm",
            "Work": "Test Label",
            "Genre": "Trance",
            "TrackNumber": 130,
            "Year": 2020,
            "MyRating": 80,
            "Location": "C:\\Music\\test_track.mp3"
        }

This one, checking an invalid Camelot Notation:

Python
@patch('pathlib.Path.exists')
    def test_invalid_album_not_camelot(self, mock_exists):
        """Test that an invalid Camelot notation fails validation."""
        # Mock file existence check
        mock_exists.return_value = True
        
        invalid_track_data = {
            "Name": "Test Track (Original Mix)",
            "Artist": "Test Artist",
            "Album": "Invalid Key",  # Not a valid Camelot notation
            "Work": "Test Label",
            "Genre": "Trance",
            "TrackNumber": 130,
            "Year": 2020,
            "MyRating": 80,
            "Location": "C:\\Music\\test_track.mp3"
        }
        
        with pytest.raises(ValueError, match="Album must be a valid Camelot notation"):
            Track(**invalid_track_data)

And this one, checking what happens with an incomplete DataFrame:

Python
    @patch('wolfie_exportvalidator_itunes.detect_file_encoding')
    @patch('pandas.read_csv')
    def test_load_itunes_data_missing_columns(self, mock_read_csv, mock_detect_encoding):
        """Test loading iTunes data with missing columns."""
        # Setup mocks
        mock_detect_encoding.return_value = 'utf-8'
        mock_df = pd.DataFrame({
            'Name': ['Test Track (Original Mix)'],
            'Artist': ['Test Artist'],
            # Missing required columns
        })
        mock_read_csv.return_value = mock_df
        
        # Call function and verify it raises an error
        with pytest.raises(ValueError, match="Missing expected columns"):
            load_itunes_data(Path('dummy_path.txt'))

I’ll include the whole test suite in my GitHub repo. Let’s conclude with pytest‘s output:

Plaintext
collected 41 items                                                                                                                                        

tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_valid_track PASSED                                                                [  2%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_valid_track_boundary_values PASSED                                                [  4%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_invalid_name_no_parentheses PASSED                                                [  7%]
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_empty_name PASSED                                                                 [  9%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_empty_artist PASSED                                                               [ 12%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_empty_work PASSED                                                                 [ 14%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_empty_genre PASSED                                                                [ 17%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_invalid_album_not_camelot PASSED                                                  [ 19%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_valid_camelot_notations PASSED                                                    [ 21%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_invalid_bpm_range_high PASSED                                                     [ 24%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_invalid_bpm_range_low PASSED                                                      [ 26%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_invalid_year_range_early PASSED                                                   [ 29%]
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_invalid_year_range_future PASSED                                                  [ 31%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_invalid_rating_not_multiple PASSED                                                [ 34%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_invalid_rating_too_low PASSED                                                     [ 36%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_invalid_rating_too_high PASSED                                                    [ 39%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_null_track_number PASSED                                                          [ 41%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_null_year PASSED                                                                  [ 43%] 
tests/test_wolfie_exportvalidator_itunes.py::TestTrackModel::test_null_rating PASSED                                                                [ 46%]
tests/test_wolfie_exportvalidator_itunes.py::TestFileOperations::test_detect_file_encoding PASSED                                                   [ 48%] 
tests/test_wolfie_exportvalidator_itunes.py::TestFileOperations::test_detect_file_encoding_latin1 PASSED                                            [ 51%] 
tests/test_wolfie_exportvalidator_itunes.py::TestFileOperations::test_detect_file_encoding_no_result PASSED                                         [ 53%] 
tests/test_wolfie_exportvalidator_itunes.py::TestFileOperations::test_load_itunes_data_success PASSED                                               [ 56%]
tests/test_wolfie_exportvalidator_itunes.py::TestFileOperations::test_load_itunes_data_missing_columns PASSED                                       [ 58%] 
tests/test_wolfie_exportvalidator_itunes.py::TestFileOperations::test_load_itunes_data_empty_dataframe PASSED                                       [ 60%] 
tests/test_wolfie_exportvalidator_itunes.py::TestValidation::test_validate_tracks_all_valid PASSED                                                  [ 63%] 
tests/test_wolfie_exportvalidator_itunes.py::TestValidation::test_validate_tracks_with_errors PASSED                                                [ 65%] 
tests/test_wolfie_exportvalidator_itunes.py::TestValidation::test_duplicate_location PASSED                                                         [ 68%] 
tests/test_wolfie_exportvalidator_itunes.py::TestValidation::test_analyze_errors PASSED                                                             [ 70%] 
tests/test_wolfie_exportvalidator_itunes.py::TestValidation::test_analyze_errors_with_general_error PASSED                                          [ 73%] 
tests/test_wolfie_exportvalidator_itunes.py::TestValidation::test_write_error_report PASSED                                                         [ 75%]
tests/test_wolfie_exportvalidator_itunes.py::TestValidation::test_process_file_with_errors PASSED                                                   [ 78%] 
tests/test_wolfie_exportvalidator_itunes.py::TestValidation::test_process_file_no_errors PASSED                                                     [ 80%]
tests/test_wolfie_exportvalidator_itunes.py::TestParamsData::test_bpm_range_valid PASSED                                                            [ 82%] 
tests/test_wolfie_exportvalidator_itunes.py::TestParamsData::test_year_range_valid PASSED                                                           [ 85%] 
tests/test_wolfie_exportvalidator_itunes.py::TestParamsData::test_rating_range_valid PASSED                                                         [ 87%] 
tests/test_wolfie_exportvalidator_itunes.py::TestParamsData::test_camelot_notations_valid PASSED                                                    [ 90%] 
tests/test_wolfie_exportvalidator_itunes.py::TestMain::test_main_with_files PASSED                                                                  [ 92%] 
tests/test_wolfie_exportvalidator_itunes.py::TestMain::test_main_no_files PASSED                                                                    [ 95%] 
tests/test_wolfie_exportvalidator_itunes.py::TestMain::test_main_with_exception PASSED                                                              [ 97%]
tests/test_wolfie_exportvalidator_itunes.py::TestMain::test_main_with_critical_exception PASSED                                                     [100%] 

=================================================================== 41 passed in 0.20s =================================================================== 

I had a very positive experience overall! Working with Amazon Q allowed me to write the tests more quickly than I could have done on my own. We would have been even faster if I had put more thought into my initial prompt. Additionally, since Q Developer offers a generous free tier, it didn’t cost me anything.

GitHub Repo

I have committed my Pydantic data validation script, test suite and documentation in the repo below:

GitHub-BannerSmall

Note that the parameters are decoupled from the Pydantic script. This will allow me to reuse some parameters across future validation scripts and has enabled me to exclude the system parameters from the repository.

Summery

In this post, I used the Pydantic Python library to create data validation and observability processes for my Project Wolfie iTunes data.

I found Pydantic very impressive! Its simplicity, functionality and interoperability make it an attractive addition to Python data pipelines, and its strong community support keeps Pydantic relevant and current. Additionally, Pydantic’s presence in FastAPI, PydanticAI and a managed AWS Lambda layer enables rapid integration and seamless deployment. I see many applications for it within Project Wolfie.

There’s lots more to Pydantic – this Pixegami video is a great walkthrough of Pydantic in action:

If this post has been useful then the button below has links for contact, socials, projects and sessions:

SharkLinkButton 1

Thanks for reading ~~^~~

Categories
Data & Analytics

SQL Workbench: On-Demand DuckDB-Wasm With No Bill

In this post, I try Tobias Müller‘s free SQL Workbench tool powered by DuckDB-Wasm. And maybe make one or two gratuitous duck puns.

Introduction

Hopefully like many people in tech, I have a collection of articles, repos, projects and whatnot saved under the banner of “That sounds cool / looks interesting / feels useful – I should check that out at some point.” Sometimes things even come off that list…

DuckDB was on my list, and went to the top when I heard about Tobias Müller’s free SQL Workbench tool powered by DuckDB-Wasm. What I heard was very impressive and pushed me to finally examine DuckDB up close.

So what is DuckDB?

About DuckDB

This section examines DuckDB and DuckDB-Wasm – core components of SQL Workbench.

DuckDB

DuckDB is an open-source SQL Online Analytical Processing (OLAP) database management system. It is intended for analytical workloads, and common use cases include in-process analytics, exploration of large datasets and machine learning model prototyping. DuckDB was released in 2019 and reached v1.0.0 on June 03 2024.

DuckDB is an embedded database that runs within a host process. It is lightweight, portable and requires no separate server. It can be embedded into applications, scripts, and notebooks without bespoke infrastructure.

DuckDB uses a columnar storage format and supports standard SQL features like complex queries, joins, aggregations and window functions. It can be used with programming languages like Python and R.

In this Python example, DuckDB:

  • Creates an in-memory database
  • Defines a users table.
  • Inserts data into users.
  • Runs a query to fetch results.
Python
import duckdb

# Create a new DuckDB database in memory
con = duckdb.connect(database=':memory:')

# Create a table and insert some data
con.execute("""
CREATE TABLE users (
    user_id INTEGER,
    user_name VARCHAR,
    age INTEGER
);
""")

con.execute("INSERT INTO users VALUES (1, 'Alice', 30), (2, 'Bob', 25), (3, 'Charlie', 35)")

# Run a query
results = con.execute("SELECT * FROM users WHERE age > 25").fetchall()
print(results)

DuckDB-Wasm

Launched in 2021, DuckDB WebAssembly (Wasm) is a version of the DuckDB database that has been compiled to run in WebAssembly. This lets DuckDB run in web browser processes and other environments with WebAssembly support.

So what’s WebAssembly? Don’t worry – I didn’t know either and so deferred to an expert:

Since all data processing happens in the browser, DuckDB-Wasm brings new benefits to DuckDB. In-browser dashboards and data analysis tools with DuckDB-Wasm enabled can operate without any server-side processing, improving their speed and security. If data is being sent somewhere then DuckDB-Wasm can handle ETL operations first, reducing both processing time and cost.

Running in-browser SQL queries also removes the need for setting up database servers and IDEs, reducing patching and maintenance (but don’t forget about the browser!) while increasing availability and convenience for remote work and educational purposes.

Like DuckDB, it is open-source and the code is on GitHub. A DuckDB Web Shell is available at shell.duckdb.org and there are also more technical details on DuckDB’s blog.

SQL Workbench

This section examines SQL Workbench and tries out some of its core features.

About SQL Workbench

Tobias Müller produced SQL Workbench in January 2024. He integrated DuckDB-Wasm into an AWS serverless static site and built an online interactive tool for querying data, showing results and generating visualizations.

SQL Workbench is available at sql-workbench.com. Note that it doesn’t currently support mobile browsers. Tobias has also written a great tutorial post that includes:

and loads more content in addition that I won’t reproduce here. So go and give Tobias some traffic!

Layout

The core SQL Workbench components are:

  • The Object Explorer section shows databases, schema and tables:
2024 06 26 SQLWorkbenchSchema
  • The Tools section shows links, settings and a drag-and-drop section for adding files (more on that later):
2024 06 26 SQLWorkbenchOptions
  • The Query Pane for writing SQL, which preloads with the below script upon each browser refresh:
2024 06 26 SQLWorkbenchQueryWindow
  • And finally, the Results Pane shows query results and visuals:
2024 06 26 SQLWorkbenchResultsWindow

Next, let’s try it out!

Sample Queries

SQL Workbench opens with a pre-loaded script that features:

  • Instructions:
SQL
-- WELCOME TO THE ONLINE SQL WORKBENCH!
-- To run a SQL query in your browser, select the query text and press:
-- CTRL + Enter (Windows/Linux) / CMD + Enter (Mac OS)

-- See https://duckdb.org/docs/sql/introduction for more info about DuckDB SQL syntax
  • Example queries using Parquet files:
SQL
-- Remote Parquet scans:
SELECT * FROM 'https://shell.duckdb.org/data/tpch/0_01/parquet/orders.parquet' LIMIT 1000;

SELECT avg(c_acctbal) FROM 'https://shell.duckdb.org/data/tpch/0_01/parquet/customer.parquet';

SELECT count(*)::int as aws_service_cnt FROM 'https://raw.githubusercontent.com/tobilg/aws-iam-data/main/data/parquet/aws_services.parquet';

SELECT * FROM 'https://raw.githubusercontent.com/tobilg/aws-edge-locations/main/data/aws-edge-locations.parquet';

SELECT cloud_provider, sum(ip_address_cnt)::int as cnt FROM 'https://raw.githubusercontent.com/tobilg/public-cloud-provider-ip-ranges/main/data/providers/all.parquet' GROUP BY cloud_provider;

SELECT * FROM 'https://raw.githubusercontent.com/tripl-ai/tpch/main/parquet/lineitem/part-0.parquet';
  • Example query using a CSV file:
SQL
-- Remote CSV scan
SELECT * FROM read_csv_auto('https://raw.githubusercontent.com/tobilg/public-cloud-provider-ip-ranges/main/data/providers/all.csv');

These queries are designed to show SQL Workbench’s capabilities and the style of SQL queries it can run (expanded on further here). They can also be used for experimentation, so here goes!

Firstly, I’ve added an ORDER BY to the first Parquet query to make sure the first 1000 orders are returned:

SQL
-- Show first 1000 orders for Parquet file
SELECT * 
FROM 'https://shell.duckdb.org/data/tpch/0_01/parquet/orders.parquet'
ORDER BY o_orderkey
LIMIT 1000;
2024 06 26 SQL WorkbenchQueryLimit1000

Secondly, I’ve used the COUNT and CAST functions to count the total orders and return the value as an integer:

SQL
-- Show number of orders for Parquet file
SELECT CAST(COUNT(o_orderkey) AS INT) AS orderkeycount 
FROM 'https://shell.duckdb.org/data/tpch/0_01/parquet/orders.parquet';
2024 06 26 SQL WorkbenchQueryOrderCount

Finally, this query captures all 1996 orders in a CTE and uses it to calculate the 1996 order price totals grouped by order priority:

SQL
-- Show 1996 order totals in priority order
WITH cte_orders1996 AS (
    SELECT o_orderkey, o_custkey, o_totalprice, o_orderdate, o_orderpriority
    FROM 'https://shell.duckdb.org/data/tpch/0_01/parquet/orders.parquet'
    WHERE YEAR(o_orderdate) = 1996
)

SELECT SUM(o_totalprice) AS totalpriorityprice, o_orderpriority 
FROM cte_orders1996 
GROUP BY o_orderpriority 
ORDER BY o_orderpriority;
2024 06 26 SQL WorkbenchQueryCTE

Now let’s try some of my data!

File Imports

SQL Workbench can also query data stored locally (as well as remote data which is out of this post’s scope). For this section I’ll be using some Parquet files generated by my WordPress bronze data orchestration process: posts.parquet and statistics_pages.parquet.

After dragging each file onto the assigned section they quickly appear as new tables:

2024 06 26 SQLWorkbenchSchemaWordPress

These tables show column names and inferred data types when expanded:

2024 06 26 SQLWorkbenchSchemaWordPressExpanded

Queries can be written manually or by right-clicking the tables in the Object Explorer. A statistics_pages SELECT * test query shows the expected results:

2024 06 26 SQLWorkbenchQueryStatsPages

As posts.parquet‘s ID primary key column matches statistics_pages.parquet‘s ID foreign key column, the two files can be joined on this column using an INNER JOIN:

SQL
SELECT 
  p.post_title,
  sp.type, 
  sp.date as viewdate,
  sp.count
FROM 
  'statistics_pages.parquet' AS sp
  INNER JOIN 'posts.parquet' AS p ON sp.id = p.id
 ORDER BY sp.date
2024 06 26 SQLWorkbenchQueryJoin

I can use this query to start making visuals!

Visuals

Next, let’s examine SQL Workbench’s visualisation capabilities. To help things along, I’ve written a new query showing the cumulative view count for my Using Athena To Query S3 Inventory Parquet Objects post (ID 92):

SQL
 SELECT 
  sp.date AS viewdate,
  SUM(sp.count) OVER (PARTITION BY sp.id ORDER BY sp.date) AS cumulative_views
FROM 
  'statistics_pages.parquet' AS sp
  INNER JOIN 'posts.parquet' AS p ON sp.id = p.id
WHERE p.id = 92
ORDER BY sp.date;
2024 06 26 SQLWorkbenchQueryCumulViews

(On reflection the join wasn’t necessary for this output, but anyway… – Ed)

Tobias has written visualisation instructions here so I’ll focus more on how my visual turns out. After the query finishes running, several visual types can be selected:

2024 06 26 SQLWorkbenchVisualizationsOptions

Selecting a visual quickly renders it in the browser. Customisation options are also available. That said, in my case the first chart was mostly fine!

2024 06 26 SQLWorkbenchVisualizationsLineChart

All I need to do is move viewdata (bottom right) to the Group By bin (top right) to add dates to the X-axis.

Exports & Cleanup

Finally, let’s examine SQL Workbench’s export options. And there’s plenty to choose from!

2024 06 26 SQLWorkbenchExports

The CSV and JSON options export the query results in their respective formats. The Arrow option exports query results in Apache Arrow format, which is designed for high-speed in-memory data processing and so compliments DuckDB very well. Dremio wrote a great Arrow technical guide for those curious.

The HTML option produces an interactive graph with tooltips and configuration options. Handy for sharing and I’m sure there’ll be a way to use the backend code for embedding if I take a closer look. The PNG has me covered in the meantime, immediately producing this image:

athenainventorycumulativeCropped
(Image cropped for visibility – Ed)

Finally there’s the Config.JSON option, which exports the visual’s settings for version control and reproducibility:

JSON
{"version":"2.10.0","plugin":"Y Line","plugin_config":{},"columns_config":{},"settings":true,"theme":"Pro Light","title":"Using Athena To Query S3 Inventory Parquet Objects Cumulative Views","group_by":["viewdate"],"split_by":[],"columns":["cumulative_views"],"filter":[],"sort":[],"expressions":{},"aggregates":{}}

All done! So how do I remove my data from SQL Workbench? I…refresh the site! Because SQL Workbench uses DuckDB, my data has only ever been stored locally and in memory. This means the data only persists until the SQL Workbench page is closed or reloaded.

Summary

In this post, I tried Tobias Müller’s free SQL Workbench tool powered by DuckDB-Wasm.

SQL Workbench and DuckDB are both very impressive! The results they produce with no servers or data movement are remarkable, and their tooling greatly simplifies working with modern data formats.

DuckDB makes lots of sense as an AWS Lambda layer, and DuckDB-Wasm simplifies in-browser analytics at a time when being data-driven and latency-averse are constantly gaining importance. I don’t feel like I’ve even scratched the surface of SQL Workbench either, so go and check it out!

Phew – no duck puns. No one has time for that quack. If this post has been useful then the button below has links for contact, socials, projects and sessions:

SharkLinkButton 1

Thanks for reading ~~^~~