London, United Kingdom

Federico Terzi

Senior Site Reliability Engineer

I run infrastructure for large products. When I do my job well, nobody notices.

Email me LinkedIn

About

I'm a Senior Site Reliability Engineer with more than ten years of experience, mostly in fintech and large messaging platforms. Day to day I look after Kubernetes clusters and the stateful systems that live on them: ScyllaDB, Kafka, PostgreSQL, and the monitoring around all of it.

I manage everything as code, usually Terraform plus automation written in Go, Python or Rust. This site is a small Rust binary running on infrastructure I terraform for fun.

Experience

  1. OneSignal

    Senior Site Reliability Engineer · London, UK

    2023 - Present

    • I help run the Kubernetes platform and the data stores behind it (ScyllaDB, Kafka, Pulsar, PostgreSQL, ClickHouse), with over 400 merged changes in the infrastructure monorepo.
    • Led the move from ZooKeeper to KRaft and took Kafka from 3.5 to 4.1 across staging and production. Also upgraded the whole PostgreSQL fleet from 15 to 17.
    • Introduced the Atlantis and Terragrunt workflow the whole org now uses for Terraform changes, and keep an eye on infrastructure costs.
    • Kubernetes
    • GCP
    • Terraform
    • ScyllaDB
    • Kafka
    • Prometheus
  2. Blockchain.com

    Senior Site Reliability Engineer · London, UK

    2022 - 2023

    • Replaced the old statsd and InfluxDB metrics stack with Prometheus and rebuilt the alerting on top of it.
    • Cut infrastructure spend by roughly 23%, partly by moving from Nomad Enterprise to the open source version.
    • Prometheus
    • Nomad
    • Kubernetes
    • AWS
  3. Freetrade

    Senior Site Reliability Engineer · London, UK

    2020 - 2022

    • Built a ChatOps service in TypeScript that handles GCP access requests for the whole engineering team.
    • Moved the existing infrastructure to code with a CircleCI pipeline and worked with the security team on VPNs and third-party connectivity.
    • GCP
    • Terraform
    • TypeScript
    • CircleCI
  4. Zopa

    Senior Site Reliability Engineer · London, UK

    2018 - 2020

    • Led the migration of a legacy Kubernetes cluster to EKS: 180 backend nodes and around 3,600 deployments.
    • Built the Concourse CI and GitHub Enterprise infrastructure on AWS with Terraform. Wrote a couple of Go services too, one for commit status checks and one that updates the status page from Prometheus alerts.
    • Kubernetes
    • AWS
    • Terraform
    • Go
    • Elasticsearch
  5. The Foundry

    DevOps Engineer · London, UK

    2017 - 2018

    • Rebuilt the CI system with Conan, CMake and Jenkins Pipelines. Product builds got up to six times faster.
    • Jenkins
    • CMake
    • VMware
  6. Netidea Webranking

    Web Analytics Specialist · Italy

    2016 - 2017

    • Analytics implementations for enterprise clients, plus Python and Flask tools to automate the repetitive parts.
    • Python
    • JavaScript
  7. CG S.r.l.

    Developer & System Administrator · Italy

    2008 - 2011

    • Wrote internal project management software in C# with MySQL and administered the office Windows and Linux machines.
    • C#
    • MySQL
    • Linux

Skills

Platform

  • Kubernetes
  • Terraform
  • GCP
  • AWS
  • Linux
  • Docker

Data & streaming

  • ScyllaDB
  • Kafka
  • Pulsar
  • PostgreSQL
  • ClickHouse
  • Elasticsearch

Observability

  • Prometheus
  • Grafana
  • Mimir
  • Alertmanager
  • SLOs

Languages

  • Go
  • Python
  • Rust
  • TypeScript
  • Bash

Education

Contact

Let’s keep things running.

Always happy to talk about infrastructure and reliability work.