Curriculum VitaeLondon, United Kingdom

Federico Terzi

Senior Site Reliability Engineer

I run infrastructure for large products. When I do my job well, nobody notices.

About

I'm a Senior Site Reliability Engineer with more than ten years of experience, mostly in fintech and large messaging platforms. Day to day I look after Kubernetes clusters and the stateful systems that live on them: ScyllaDB, Kafka, PostgreSQL, and the monitoring around all of it.

I manage everything as code, usually Terraform plus automation written in Go, Python or Rust. This site is a small Rust binary running on infrastructure I terraform for fun.

Experience

2023–Present

London, UK

OneSignal Senior Site Reliability Engineer

  • I help run the Kubernetes platform and the data stores behind it (ScyllaDB, Kafka, Pulsar, PostgreSQL, ClickHouse), with over 400 merged changes in the infrastructure monorepo.
  • Led the move from ZooKeeper to KRaft and took Kafka from 3.5 to 4.1 across staging and production. Also upgraded the whole PostgreSQL fleet from 15 to 17.
  • Introduced the Atlantis and Terragrunt workflow the whole org now uses for Terraform changes, and keep an eye on infrastructure costs.

KubernetesGCPTerraformScyllaDBKafkaPrometheus

2022–2023

London, UK

Blockchain.com Senior Site Reliability Engineer

  • Replaced the old statsd and InfluxDB metrics stack with Prometheus and rebuilt the alerting on top of it.
  • Cut infrastructure spend by roughly 23%, partly by moving from Nomad Enterprise to the open source version.

PrometheusNomadKubernetesAWS

2020–2022

London, UK

Freetrade Senior Site Reliability Engineer

  • Built a ChatOps service in TypeScript that handles GCP access requests for the whole engineering team.
  • Moved the existing infrastructure to code with a CircleCI pipeline and worked with the security team on VPNs and third-party connectivity.

GCPTerraformTypeScriptCircleCI

2018–2020

London, UK

Zopa Senior Site Reliability Engineer

  • Led the migration of a legacy Kubernetes cluster to EKS: 180 backend nodes and around 3,600 deployments.
  • Built the Concourse CI and GitHub Enterprise infrastructure on AWS with Terraform. Wrote a couple of Go services too, one for commit status checks and one that updates the status page from Prometheus alerts.

KubernetesAWSTerraformGoElasticsearch

2017–2018

London, UK

The Foundry DevOps Engineer

  • Rebuilt the CI system with Conan, CMake and Jenkins Pipelines. Product builds got up to six times faster.

JenkinsCMakeVMware

2016–2017

Italy

Netidea Webranking Web Analytics Specialist

  • Analytics implementations for enterprise clients, plus Python and Flask tools to automate the repetitive parts.

PythonJavaScript

2008–2011

Italy

CG S.r.l. Developer & System Administrator

  • Wrote internal project management software in C# with MySQL and administered the office Windows and Linux machines.

C#MySQLLinux

Skills

Platform

KubernetesTerraformGCPAWSLinuxDocker

Data & streaming

ScyllaDBKafkaPulsarPostgreSQLClickHouseElasticsearch

Observability

PrometheusGrafanaMimirAlertmanagerSLOs

Languages

GoPythonRustTypeScriptBash

Education

2015 - 2016

Bologna Business School, Università di Bologna

Master in Data Science

2008 - 2015

Università di Modena e Reggio Emilia

BSc Computer Science

Contact

Write to [email protected], or find me on LinkedIn and GitHub.

London, United Kingdom