Several years of hands-on experience as an SRE, Storage Engineer or Platform Engineer running production storage infrastructure, with solid experience deploying, operating and debugging Ceph in production.
You know the Ceph fundamentals from real-world operations: CRUSH, pools and placement groups, replication vs. erasure coding, BlueStore, scrubbing, recovery and backfill under load.
You have managed storage infrastructure end to end, from the physical layer with drives, firmware, BIOS, controllers and hardware diagnostics through rebalancing, host drains, capacity expansion and hardware generation migrations.
You are comfortable using Kubernetes as a control plane, not only as a workload runtime, and understand operators/controllers, custom resources, reconciliation and GitOps. You can also move confidently around OpenStack and debug issues involving Ironic, Neutron, Nova or Glance when needed.
You have dealt with situations where durability really matters – such as degraded or near-full clusters, recovery under load or split-brain conditions – and you are used to making changes in a staged, evidence-based and reversible way.
You have strong Linux and bare-metal skills, understand the block layer, filesystems and I/O behaviour, and are comfortable working with Ansible, Terraform and GitOps.
AI-assisted engineering is already part of your day-to-day work. You have practical experience with LLMs and agentic tools and know where they can meaningfully support development, testing, reviews or operations.
Ideally, you also bring experience with RGW / S3, RBD mirroring, CephFS or ceph-csi, deeper OpenStack storage integrations and Go and/or Python. Experience around VLAN/BGP, storage performance tuning, NVMe/BlueStore, observability, data protection, encryption, auto-remediation, security baselines or multi-site storage is also relevant for the role.
You work autonomously, bring a strong sense of ownership, and are comfortable debugging problems where the actual root cause may sit somewhere between storage, networking, Kubernetes and OpenStack. You would rather verify what is happening from evidence than assume how the system should behave.
We work in an international environment in English, so you should feel comfortable discussing technical topics, documenting decisions and working with the team in English.