David Collom

Building Scalable, Secure and High-Performant Platforms with Kubernetes
Cloud-Native Specialist | Homelabber at Heart | Always Learning & Building

[email protected] // Co Durham, UK


// About Me

Platform engineer and Solutions Architect with a software engineering background and more than eight years working with Kubernetes across cloud and on-premises environments. I focus on infrastructure automation, platform design, Terraform, developer experience and the operational work that helps teams ship reliably. I have spoken at KubeCon and ArgoCon on Kubernetes cost optimisation and reliability, and I contribute to open-source platform and operations tools. My strongest work sits where hands-on engineering, technical leadership and platform outcomes meet.

// Skills

  • Languages / DBs

  • Go, Python, Ruby, Redis, MySQL, PostgreSQL
  • Frameworks

  • Controller Runtime, OpenAPI / Swagger, Rails, Grape, KOPF, Sinatra
  • Applications / Toolkits

  • Kubernetes, Terraform, Docker, Chef, Git, Terragrunt, Nginx, Apache, HAProxy, Argo Project, Argo CD, Flux, Ansible, HashiCorp Vault, Consul
  • Observability

  • Prometheus, OpenTelemetry, Grafana, InfluxDB, Thanos, VictoriaMetrics
  • Cloud Services / Providers

  • Google Cloud Platform (GCP), Digital Ocean, Amazon AWS
  • Other

  • Team Leadership, Mentoring and Coaching, Open-source Contributions, Issue Triage and Planning, System Design and Architecture, Pragmatic Problem Solving

// Experiences

Solutions Architect

 
Komodor  (Remote (UK), EMEA)
  January 2026 - Present
  • Lead post-sales architecture for EMEA customers running Komodor in production Kubernetes environments. The work covers discovery, onboarding, adoption reviews and technical escalations.
  • Built komodor-klaudia-sync to keep Klaudia knowledge and blueprints in version control and synchronise them from GitHub Actions. Its CI runs formatting, go vet, govulncheck and race tests, while GoReleaser publishes binaries and signed containers to GHCR.
  • Built komodor-security-reporter, a Kubernetes-native Go service that resolves running workload images to digests, normalises scanner findings and sends deduplicated, low-noise events to Komodor. It runs with minimal Kubernetes privileges and has race/coverage tests, linting, benchmarks and multi-platform releases.
  • When a SIEM requirement blocked a customer onboarding and rollout, built a private Go bridge around the Komodor API. It included unit and integration tests, benchmarks, CI/CD and a repeatable release process.
  • Turned deployment gaps found in customer environments into Terraform provider work and merged public Helm chart changes covering RBAC, security contexts, custom volumes, custom CAs and chart hardening.
  • Built and reviewed public integration prototypes for gaps in the main platform, including Prometheus/Alertmanager event ingestion and provider-neutral RCA-to-Slack deployments for Google Cloud Run and AWS Lambda.
  • Contributed to Komodor's closed-source core product in the first week, fixing public API service issues and adding API capability required by a customer.
  • Work with Product, Engineering and GTM to turn recurring customer constraints into clear requirements and practical next steps.

Staff Solutions Engineer

 
Jetstack / Venafi / CyberArk  (Remote (UK))
  November 2020 - January 2026
  • Delivered hands-on customer engagements across Kubernetes, cloud-native security, platform architecture, GitOps, infrastructure as code, automation and application integration.
  • Converted customer outcomes into scoped Statements of Work, delivery estimates, milestones, dependencies, skills requirements and implementation plans.
  • Led distributed project teams while staying close to the engineering work: architecture, implementation, technical governance, customer communication and problem resolution.
  • Built Kubernetes platforms, bespoke integrations, controllers and operators that helped customers connect cloud-native tools with existing enterprise systems.
  • Helped organisations move from manual infrastructure processes towards GitOps, infrastructure as code and Kubernetes operating models that platform teams could maintain.
  • Built repeatable training and workshop platforms that gave participants hands-on cloud-native environments with much less setup work for instructors and attendees.
  • Contributed to open-source Kubernetes and platform projects, including operator, controller, security and resource-optimisation work.
  • Mentored engineers between customer engagements, helped shape delivery approaches and fed practical customer lessons back to product and engineering teams.
  • Spoke at KubeCon + CloudNativeCon North America 2023 and ArgoCon Europe 2023, using customer and engineering experience to explain Kubernetes cost, reliability and platform practices.

Senior Infrastructure Engineer

 
Credit Karma  (London, UK)
  July 2019 - October 2020
  • Managed production Kubernetes platforms and shared infrastructure services for engineering teams across cloud and on-premises environments.
  • Used Terraform and Terragrunt to make infrastructure changes more repeatable, reviewable and easier to operate.
  • Worked on capacity planning and predictive alerting so platform pressure could be spotted before it became an incident.
  • Managed service-mesh infrastructure and supported migration work across Linkerd and Envoy-based environments.
  • Contributed to the launch of an independent Google Cloud Platform environment during a wider organisational separation.
  • Upstreamed security-driven maintenance to Kubernetes ExternalDNS and Prometheus Blackbox Exporter, and built controller-runtime automation for service-account-specific registry credentials.

Infrastructure Platform Engineer (Kubernetes) / Senior DevOps Engineer

 
Sky Betting and Gaming  (Leeds, UK)
  June 2017 - July 2019
  • Owned and improved on-premises and AWS-hosted Kubernetes platforms used by internal engineering teams.
  • Developed custom Kubernetes cloud-controller capability for on-premises node provisioning, bringing cloud-provider style automation into private infrastructure.
  • Designed availability-zone patterns and automated service-health checks to improve resilience, fault tolerance and incident detection.
  • Standardised deployment and configuration across hybrid infrastructure with Terraform.
  • Improved automated patching with Jenkins, Chef, Bash, Python and Ruby, covering alert silencing, service draining, host rebooting and safe restoration.
  • Reduced toil and manual error by automating maintenance workflows that crossed application, monitoring and infrastructure boundaries.
  • Worked with SRE, performance-testing and application teams so new services could join shared operational processes safely.
  • Built Go-based tooling and Kubernetes automation for practical platform engineering needs, including inspection, validation and operational workflows.

Senior Software Developer

 
William Hill  (Leeds, UK)
  February 2016 - June 2017
  • Acted as Developer Lead and Senior Developer for WH Cloud, including running the project independently for several months before leaving the organisation.
  • Led or substantially contributed to moving a monolithic application and API towards service-oriented architecture and distributed systems.
  • Used that decoupling work to speed up proof-of-concept activity with third-party cloud providers and give teams a clearer path to microservices adoption.
  • Evaluated Kubernetes, Mesos/Marathon, Docker Swarm, AWS, Azure and VMware cloud platforms, comparing the architecture and operating trade-offs rather than chasing a single tool.
  • Built internal tooling, APIs and platform workflows that supported a new datacentre design and made application onboarding more self-service for development teams.
  • Designed and delivered workshops, training environments and pre-provisioned laptops so developers could learn and adopt Puppet, configuration-management practices and the WH Cloud platform.
  • Developed software to visualise and support firewall automation, improving approval flows between Engineering, Networks and Information Security teams.
  • Planned engineering workloads, mentored developers and systems engineers, joined interviews and supported performance-management activities.
  • Improved delivery practices through peer review, more unit testing, benchmarking and Software Development Life Cycle governance.

Software Developer

 
William Hill  (Leeds, UK)
  May 2014 - February 2016
  • Worked on WH Cloud, an internal platform that provided customised Platform as a Service capabilities on existing on-premises and VMware vSphere infrastructure.
  • Built an automated containerised infrastructure product with Mesos and Marathon, including bespoke integrations into existing enterprise systems and VMware vSphere provisioning.
  • Contributed to platform work that reduced delivery or scaling of business applications from approximately three to six weeks to around one to two hours.
  • Developed automation, APIs and integration capabilities that made infrastructure delivery repeatable, self-service and easier for development teams to consume.
  • Designed and delivered internal technical training and onboarding workshops to support adoption of the platform, Puppet and container-based operating practices.

Software Developer

 
City Electrical Factors  (Durham, UK)
  June 2012 - May 2014
  • Developed and maintained bespoke e-commerce, content-management, CRM, ERP and reporting systems for customers, branches, warehouses and internal business teams.
  • Combined Ruby on Rails and PHP development with infrastructure architecture, deployment automation and production support.
  • Helped move hosting away from a single physical appliance towards distributed virtualised infrastructure with load balancing, replication, clustering and automated provisioning.
  • Introduced Puppet OSS, Vagrant, Docker and Dokku-based workflows to improve local development, repeatable provisioning and application deployment.
  • Supported daily releases through code review, deployment help and close work with other software engineers.
  • Mentored junior and mid-level engineers and advised senior stakeholders on architecture, cost and infrastructure direction.
  • Operated integrations between public web systems and internal business services, including ActiveMQ messaging and early REST API gateway work.

Developer & System Administrator

 
Visualsoft eCommerce  (Stockton-on-Tees, UK)
  August 2008 - June 2012
  • Managed Linux servers and production infrastructure hosting more than 400 e-commerce websites.
  • Built, patched, secured and supported Red Hat Enterprise Linux and Apache/PHP systems, including monitoring, backups, DNS and out-of-hours incident response.
  • Supported PCI-compliant operations through vulnerability remediation, access controls, log auditing, file permissions and audit documentation.
  • Diagnosed availability and performance incidents across a large multi-customer estate with limited monitoring and instrumentation.
  • Developed and supported internal systems and customer-facing platform features alongside production systems administration work.
  • Implemented and supported eBay and Amazon marketplace synchronisation, dynamic price matching and ePOS/till integrations for products, inventory, orders and sales data.
  • Helped the platform move from physical Rackspace-hosted infrastructure towards a VMware/vSphere-based estate while keeping customer-facing services running.

// Achievements & certifications

Speaker: Fusion Meetup - Cost Confessions

  2024  
View
 
Watch

On 12 June 2024, Becky Pauley and I presented "Cost Confessions: Tales of Overspending and Redemption" at Fusion Meetup.

The session offered a vendor-agnostic approach to controlling Kubernetes cloud spend. We covered cost visibility, resource right-sizing, appropriate use of discounts and the practical pitfalls that can turn an apparently efficient platform into an expensive one.

Bringing the talk to a local engineering community created space for direct questions and discussion. It was a useful reminder that meetups are a feedback loop between shared experience and day-to-day practice.


Speaker: Yorkshire DevOps - GKE Overview and History

  2023  
Slides

I presented at Yorkshire DevOps to an audience of over 60 attendees, covering the evolution and features of Google Kubernetes Engine (GKE). The session covered GKE's history and capabilities, including AutoPilot, Release Channels, and add-ons such as Backups, Istio, Config-Connector, Ingress Controller, Secrets, and IAM Integration.

The presentation gave attendees a practical overview of GKE's strengths and its role in cloud-native architecture. The discussion with attendees centred on Kubernetes and cloud infrastructure management in real engineering teams.


Speaker: KubeCon + CloudNativeCon North America 2023

  2023  
View
 
Watch

I presented alongside a colleague at KubeCon NA on "Kubernetes Confessions: Tales of Overspending and Redemption", covering common problems in cloud cost management. Our session attracted over 360 RSVPs, including attendees from notable organisations such as Microsoft, Google, Amazon, and HashiCorp. The session showed that Kubernetes cost management had become a shared concern across the Kubernetes community.

Our talk used a vendor-agnostic approach to managing Kubernetes cloud spend, with practical examples and open-source tools. Topics included visibility into cloud and Kubernetes spend, right-sizing resources, using discounts appropriately, and avoiding common pitfalls.

Attendee questions focused on practical ways to improve cloud infrastructure management without turning cost work into a blame exercise.


Speaker: CNCF-Hosted Co-Located Events Europe 2023: Argo Con

  2023  
View
 
Watch

In 2023, I delivered my first large conference talk, "How to Avoid a Kubernetes Doom Loop," at ArgoCon Europe. This lightning talk recounted an incident involving a Kubernetes cluster managing 16K Argo Workflows across 165 nodes, showing the risks of automation misconfiguration and the need for Cloud Reliability Engineering (CRE) practices.

The talk attracted over 280 RSVPs, showing strong interest in the topic within the Kubernetes community. The lesson was straightforward: proactive maintenance, bounded automation and tested recovery paths matter when cloud-native systems fail.


CS-169.1x: Software as a Service

  2013  
View