Quick Overview
Job Description
ABOUT THE TEAM & ROLE
Our SRE / Production Engineering team is responsible for keeping Talon.One reliable, scalable, and easy to operate. We work closely with engineering teams across R&D to improve how we monitor, release, troubleshoot, and run our production systems.
This is a hands-on role for someone who loves solving production-level problems, automating repetitive work, and building pragmatic tooling to make life safer and easier for the engineers around them.
ONCE YOU ARE HERE, YOU WILL:
- Eliminate Toil: Identify manual or repetitive operational friction across R&D and build clean scripts, automation, and internal tools to solve it permanently.
- Pioneer AI-Driven Operations: Design, build, and integrate AI agents to streamline operational workflows, ensuring proper guardrails, monitoring, and human oversight for safe execution.
- Level Up Incident Management: Own and optimize our Incident.io workflows, automation, and integrations. Stay closely engaged with incident response and participate in post-incident reviews to identify friction and turn learnings into improvements to tooling, coordination, and processes, without taking on incident responder responsibilities.
- Enhance Observability & System Health: Maintain and refine monitoring, alerting, and dashboards across our observability stack. You will dive into logs, metrics, and production data to investigate operational edge cases.
- Optimize Workflows & Runbooks: Partner directly with SRE and R&D teams to identify operational pain points, turning messy procedures into clear, automated runbooks.
- Support Core Production Systems: Collaborate with SREs on database maintenance tasks, health checks, and release/deployment workflows where production reliability is impacted.
WHAT WE NEED YOU TO BRING TO THE TABLE:
- 2–4 years of experience in TechOps, DevOps, SRE, Production Engineering, or a similar technical role
- Experience working with production systems in a SaaS or cloud environment
- Comfortable working with Linux, command-line tools, logs, and monitoring
- Experience with scripting or automation and a mindset of "if we do it twice, can we automate it?"
- A structured approach to troubleshooting and solving operational problems
- Proactive attitude towards improving systems, processes, and tooling
- Ability to work collaboratively with engineers across different teams
- Willingness to learn and build deeper expertise in production systems and reliability
NICE TO HAVE
- Familiarity with core SRE concepts, such as Service Level Indicators/Objectives (SLIs/SLOs) and error budgets.
- Experience with observability platforms such as Grafana, Datadog, Prometheus, or Sentry
- Experience with Kubernetes and GCP
- Experience working with APIs, integrations, CI/CD, or infrastructure automatio
OUR TECH STACK
- Cloud & Infrastructure: Google Cloud Platform (GCP), Kubernetes, containerized workloads
- Infrastructure as Code: Terraform / Helm
- Observability & Incident Ops: Grafana, Datadog, Sentry, Incident.io
- Databases: PostgreSQL
- Automation & CI/CD: Go, Python, Bash, GitHub Actions / CI/CD tooling
- 120+ team of engineers, product managers and product designers in Berlin
- Leaders with 8+ years of experience building our promotions engine
- €1,000 annual learning budget and free German language courses to boost your skills
- 30 days of annual leave, plus extra paid days for your birthday and moving day
- Home office setup budget, a monthly home office allowance
- Freedom to work from abroad for up to 90 days worldwide!
- Mental health support with nilo.health and a discounted Urban Sports Club membership
- 20% company subsidy on your pension contributions
- Subsidised BVG public transport ticket and a dog-friendly Berlin office where your furry friend is welcome
- Lease your ideal bike through BusinessBike
WHY YOU SHOULD WORK FOR US:
- The right attitude: modern methods and a diverse, creative workspace with an open and international culture
- Everyone for the product: Together we create a flexible, highly scalable product with state-of-the-art technologies. We can only succeed if everyone works as a team
- Healthy Growth: Growing our company means growing everyone in the team. We love to share knowledge and learn
- A great environment: Flexible and family-friendly environment, bright and easily accessible offices, modern software and hardware
- High flexibility degree: Prefer to work early or late at night? Do you have to pick up your children from kindergarten? Do you prefer working abroad? We believe in results and motivated employees
Do you want this job?
We’d love to hear from you! Apply directly via the form below.
Talon.One is an Equal Employment Opportunity employer that proudly pursues and hires a diverse workforce. We do not make employment decisions on the basis of race, color, religious belief, ethnic origin, nationality, sex, gender identity, sexual orientation, disability, age, military or veteran status, or any other basis protected by applicable local, state, or federal laws or prohibited by company policy. As an employer we strive for a healthy and safe workplace and strictly prohibit harassment of any kind.
Find out more about our Candidate Privacy Policy.
Similar jobs
- EA
(Senior) DevOps & QA Engineer (d/m/w) - AI-based Autonomy Hub
NewE Airbus Defence and Space GmbH
Stuttgart, Baden-Württemberg🇩🇪Hybrid1 hour agoDockerAWSAnsible+11Technology - EA
Junior Cloud & DevOps Engineer / Kubernetes Specialist (f/m/d)
NewE Airbus Defence and Space GmbH
München, Bayern🇩🇪Hybrid1 hour agoDockerNode.jsArgoCD+5Technology - A1
(Senior) Software Engineer: Product & DevOps (m/f/d)
NewA11
Berlin🇩🇪Hybrid7 hours agoTypeScriptTechnology - IC
Senior DevOps / Platform Engineer, ISR
NewIceye
Berlin🇩🇪Hybrid14 hours agoAgileKubernetesPostgreSQL+4Technology - WA
Site Reliability Engineer, Vehicle SW
NewWayve
Leonberg🇩🇪Hybrid11 hours agoRustSplunkContinuous Improvement+7Technology - DC
DevOps Engineer (x/f/m)
NewDKB Code Factory
Berlin🇩🇪Hybrid13 hours agoAWSAgileGitLab CI+5Technology