David Collom

Building Scalable, Secure and High-Performant Platforms with Kubernetes
Cloud-Native Specialist | Homelabber at Heart | Always Learning & Building

[email protected] // Co Durham, UK


// About Me

Experienced Platform Engineer with over a decade of experience in infrastructure automation and platform design, including 8+ years of deep expertise in architecting and operating Kubernetes across cloud and on-premises environments. Focused on infrastructure as code at scale, using Terraform to build robust environments that accelerate delivery and reduce drift. Demonstrated thought leadership through presentations at industry conferences (e.g., KubeCon, ArgoCon), sharing insights on Kubernetes cost optimisation and platform reliability, and actively contributed to open-source community tools. Eager to leverage unique technical skills and leadership abilities to drive success in future platforms, ensuring high-quality services and rapid delivery in a fast-paced environment.

// Skills

  • Languages / DBs

  • Go, Python, Ruby, Redis, MySQL, PostgreSQL
  • Frameworks

  • Controller Runtime, OpenAPI / Swagger, Rails, Grape, KOPF, Sinatra
  • Applications / Toolkits

  • Kubernetes, Terraform, Docker, Chef, Git, Terragrunt, Nginx, Apache, HAProxy, Argo Project, Argo CD, Flux, Ansible, HashiCorp Vault, Consul
  • Observability

  • Prometheus, OpenTelemetry, Grafana, InfluxDB, Thanos, VictoriaMetrics
  • Cloud Services / Providers

  • Google Cloud Platform (GCP), Digital Ocean, Amazon AWS
  • Other

  • Team Leadership, Mentoring and Coaching, Open-source Contributions, Issue Triage and Planning, System Design and Architecture, Pragmatic Problem Solving

// Experiences

Solutions Architect

 
Komodor  (Remote (UK), EMEA)
  January 2026 - Present
  • Lead post-sales technical discovery and solution design with Platform Engineering, SRE, infrastructure and engineering leadership teams operating complex Kubernetes environments across EMEA.
  • Translate customer platform requirements into implementation plans, adoption paths and enablement work that support successful onboarding and platform maturity.
  • Connect Komodor's operational intelligence and AI-assisted troubleshooting capabilities to customer reliability, Kubernetes operations and Developer Experience goals.
  • Build trusted relationships with technical practitioners and executive stakeholders, translating complex Kubernetes challenges into clear architectures, practical recommendations and measurable outcomes.
  • Develop customer-facing proofs of concept and API-led integrations in Go to unblock adoption and demonstrate integration patterns across REST APIs, SIEM workflows and AI agents using MCP.
  • Support Argo CD, GitOps, Terraform and Helm-based deployment patterns, including contributing field-led improvements to Komodor's public Helm charts and Terraform provider.
  • Perform proactive customer health checks, lead customer review meetings and provide technical support for escalations, renewals and expansion opportunities.
  • Capture feature requests, product gaps and recurring operational needs, feeding field insight back to Product and go-to-market teams.

Staff Solutions Engineer

 
Jetstack / Venafi / CyberArk  (Remote (UK))
  November 2020 - January 2026
  • Led hands-on enterprise consulting engagements across Kubernetes, cloud-native security, platform architecture, GitOps, infrastructure as code, automation and application integration, primarily for EMEA customers.
  • Owned technical discovery and engagement planning, converting customer outcomes into scoped Statements of Work, delivery estimates, milestones, dependencies, skills requirements and implementation plans.
  • Acted as lead technical contributor across distributed project teams, balancing architecture, direct engineering, technical governance and stakeholder communication while resolving complex delivery challenges.
  • Designed and implemented Kubernetes platforms, bespoke integrations, controllers and operators that connected cloud-native capabilities with existing enterprise systems and reduced repetitive operational effort.
  • Guided customers from manual infrastructure processes towards Kubernetes-native operating models, GitOps, infrastructure as code and sustainable platform engineering practices.
  • Built trusted relationships with engineers, programme teams, Heads of Engineering, CTOs and other senior stakeholders, communicating technical decisions, progress, risks and architectural trade-offs.
  • Architected repeatable internal and external training platforms that provisioned hands-on cloud-native environments, reduced workshop setup toil and improved the experience for instructors and participants.
  • Led distributed technical delivery across time zones, mentored engineers, contributed to open-source projects and fed customer insight back to product and engineering teams.
  • Presented at KubeCon + CloudNativeCon North America 2023 and ArgoCon Europe 2023, and represented the organisation through technical demonstrations at CNCF events.

Senior Infrastructure Engineer

 
Credit Karma  (London, UK)
  July 2019 - October 2020
  • Managed, maintained and enhanced production Kubernetes platforms and core infrastructure services supporting engineering teams across cloud and on-premises environments.
  • Automated infrastructure provisioning and change management using Terraform and Terragrunt, improving repeatability and operational control.
  • Designed capacity-planning strategies and predictive alerting to improve platform reliability, operational readiness and informed scaling decisions.
  • Managed service-mesh infrastructure and led migration work across Linkerd and Envoy-based environments.
  • Contributed to the launch of an independent platform on Google Cloud Platform, supporting organisational separation and cloud-native platform delivery.

Infrastructure Platform Engineer (Kubernetes) / Senior DevOps Engineer

 
Sky Betting and Gaming  (Leeds, UK)
  June 2017 - July 2019
  • Owned and improved multiple on-premises and AWS-hosted Kubernetes platforms, supporting platform reliability and the internal engineering teams consuming them.
  • Developed a custom Kubernetes Cloud Controller for on-premises node provisioning, extending cloud-provider automation patterns into private infrastructure.
  • Designed on-premises availability-zone capabilities and automated service-health checks to improve resilience, fault tolerance and incident detection.
  • Standardised infrastructure deployment and configuration across hybrid environments using Terraform.
  • Enhanced automated patching through Jenkins, Chef and custom Bash, Python and Ruby tooling, coordinating alert silencing, service draining, host rebooting and safe restoration.
  • Reduced operational toil and human-error risk by automating maintenance workflows that crossed application, monitoring and infrastructure boundaries.
  • Partnered with SRE, performance-testing and application teams to introduce services safely into shared operational processes and improve Kubernetes adoption.
  • Developed Go-based internal tooling and Kubernetes-native automation to address practical platform engineering requirements.

Senior Software Developer

 
William Hill  (Leeds, UK)
  February 2016 - June 2017
  • Led or substantially contributed to decomposing a monolithic application and API towards service-oriented and microservices architectures, accelerating cloud-provider proofs of concept.
  • Led technical evaluations and proofs of concept across Kubernetes, Mesos/Marathon, Docker Swarm, AWS, Azure and VMware cloud platforms, assessing architectural and operational trade-offs.
  • Developed software to visualise and support firewall automation, improving approval workflows between Engineering, Networks and Information Security teams.
  • Managed and planned engineering workloads, mentored developers and systems engineers, participated in recruitment and supported performance-management activities.
  • Improved engineering practices through peer review, increased unit-test adoption, benchmarking and Software Development Life Cycle governance.

Software Developer

 
William Hill  (Leeds, UK)
  May 2014 - February 2016
  • Worked on the WH Cloud internal platform, delivering a customised Platform as a Service on existing on-premises and VMware vSphere infrastructure.
  • Built an automated containerised infrastructure product using Mesos and Marathon with bespoke integrations into existing enterprise systems.
  • Contributed as a key team member to reducing delivery or scaling of key business applications from approximately three to six weeks to approximately one to two hours.
  • Developed automation and integration capabilities that enabled repeatable infrastructure delivery and improved developer access to platform services.
  • Designed and delivered internal technical training that supported adoption of the platform and container-based operating practices.

Software Developer

 
City Electrical Factors  (Durham, UK)
  June 2012 - May 2014
  • Developed and maintained bespoke e-commerce, content management, CRM, ERP and reporting systems supporting customer, branch, warehouse and business operations.
  • Combined Ruby on Rails and PHP software development with infrastructure architecture, deployment automation and production support.
  • Helped move hosting away from reliance on a single physical appliance towards a distributed virtualised architecture with load balancing, replication, clustering and automated provisioning.
  • Introduced and supported Puppet OSS, Vagrant, Docker and Dokku-based workflows to improve local development, repeatable provisioning and application deployment.
  • Supported a daily release cadence through code review, deployment assistance and close collaboration with software engineers.
  • Mentored junior and mid-level engineers and advised senior technical stakeholders and budget holders on architecture, cost and infrastructure strategy.
  • Operated integrations between public web systems and internal business services, including ActiveMQ-based messaging and early REST API gateway work.

Developer & System Administrator

 
Visualsoft eCommerce  (Stockton-on-Tees, UK)
  August 2008 - June 2012
  • Managed Linux servers and production infrastructure hosting more than 400 high-traffic e-commerce websites.
  • Built, patched, secured and supported Red Hat Enterprise Linux and Apache/PHP systems, including monitoring, backups, DNS and out-of-hours incident response.
  • Supported PCI-compliant operations through vulnerability remediation, access controls, log auditing, system permissions and audit documentation.
  • Diagnosed website availability and performance incidents in a large multi-customer environment with limited monitoring and instrumentation.
  • Developed and supported internal systems and customer-facing platform features alongside production systems administration responsibilities.
  • Implemented and supported eBay and Amazon marketplace synchronisation, dynamic price-matching and ePOS/till integrations for products, inventory, orders and sales data.
  • Supported the transition from physical Rackspace-hosted infrastructure towards a VMware/vSphere-based estate while maintaining customer-facing service continuity.

// Achievements & certifications

Speaker: Fusion Meetup - Cost Confessions

  2024  
View
 
Watch

On 12 June 2024, Becky Pauley and I presented “Cost Confessions: Tales of Overspending and Redemption” at Fusion Meetup.

The session offered a vendor-agnostic approach to controlling Kubernetes cloud spend. We covered cost visibility, resource right-sizing, appropriate use of discounts and the practical pitfalls that can turn an apparently efficient platform into an expensive one.

Bringing the talk to a local engineering community created space for direct questions and discussion, reinforcing the value of meetups as a feedback loop between shared experience and day-to-day practice.


Speaker: Yorkshire DevOps - GKE Overview and History

  2023  
Slides

I presented at Yorkshire DevOps to an audience of over 60 attendees, focusing on the evolution and features of Google Kubernetes Engine (GKE). The session covered GKE's history and highlighted its powerful capabilities, including AutoPilot, Release Channels, and essential add-ons such as Backups, Istio, Config-Connector, Ingress Controller, Secrets, and IAM Integration.

The presentation aimed to provide a comprehensive overview of GKE's strengths and its role in modern cloud-native architectures. Engaging with attendees at Yorkshire DevOps reinforced my commitment to sharing knowledge and promoting best practices in Kubernetes and cloud infrastructure management.


Speaker: KubeCon + CloudNativeCon North America 2023

  2023  
View
 
Watch

I presented alongside a colleague at KubeCon NA on “Kubernetes Confessions: Tales of Overspending and Redemption”, addressing critical challenges in cloud cost management. Our session attracted over 360 RSVPS, including attendees from notable organisations such as Microsoft, Google, Amazon, and HashiCorp. This experience highlighted our ability to provide actionable insights and engage a diverse audience within the Kubernetes community.

Our talk provided a vendor-agnostic approach to managing Kubernetes cloud spend efficiently, emphasising best practices and open-source tools. Key topics included gaining visibility into cloud and Kubernetes expenditures, right-sizing resources, leveraging discounts, and preemptively avoiding common pitfalls.

The positive reception from attendees reinforced our reputation as knowledgeable speakers in the Kubernetes ecosystem, We are known for delivering practical solutions and fostering continuous improvement within cloud infrastructure management.


Speaker: CNCF-Hosted Co-Located Events Europe 2023: Argo Con

  2023  
View
 
Watch

In 2023, I had the privilege of delivering my inaugural large-scale conference talk, “How to Avoid a Kubernetes Doom Loop,” at a prominent tech event. This lightning talk recounted a critical incident involving a Kubernetes cluster managing 16K Argo Workflows across 165 nodes, highlighting the risks of automation misconfiguration and emphasising the criticality of adhering to Cloud Reliability Engineering (CRE) best practices.

The talk attracted over 280 RSVPS, indicating substantial interest in the topic within the Kubernetes community and marking a successful debut on the conference speaking circuit. This experience underscored the importance of proactive maintenance and adherence to best practices in ensuring the stability and reliability of cloud-native infrastructures.


CS-169.1x: Software as a Service

  2013  
View