About us
ATOSS Software SE is one of Germany’s most successful tech growth stories. As the market leader in Workforce Management Software, we help companies work more intelligently, creatively, and humanely optimizing the balance between profitability and people.
We’re a rare company: according to Handelsblatt (10/24), just 309 public companies worldwide achieved over 20% return on sales for ten consecutive years. Only one based in Germany: ATOSS Software SE.
With 20 years of record breaking growth, over €2 billion market cap, and listings in SDAX and TecDAX, we’re scaling globally and we’re growing.
If you’re ready to drive impact in a high-performing B2B SaaS environment, this is your chance to elevate your career.
The Person You are
At ATOSS, we hire for both character and skill, seeking individuals who embody resilience, a pioneering spirit, and the passion to grow.
We value those who:
Think like entrepreneurs – taking ownership, pushing boundaries, and driving impact.
Challenge the status quo – bringing fresh ideas and bold execution to the table.
Thrive in change – seeing growth as a lifelong journey, both professionally and personally.
We don’t just equip you for work—we prepare you for life.
The Role
As a (Senior) Cloud Ops Engineer in our Cloud Operations Services (COS) department, you keep the services running that the ATOSS Cloud Solution is built on, across Azure and T Cloud Public. Some of these services are managed cloud services, others we host and operate ourselves, and you take care of both. The role has a strong DevOps character: you build CI/CD pipelines in Azure DevOps, write Infrastructure as Code, script away repetitive work, and help us grow our AI infrastructure and AI operations. You work closely with R&D, Cloud Architecture, Security, and the other COS engineers.
Key Responsibilities
Potential Areas of Ownership
Depending on team priorities, experience, and the agreed assignment, your responsibilities may include the following areas:
Cloud service operations
Deploy, operate, and maintain a broad range of selected Azure services and their counterparts on T Cloud Public, from managed platform services to compute, storage, and networking.
Run the self-hosted services within your assigned scope—such as Kafka, Keycloak, Redis, and PostgreSQL—on Kubernetes (AKS) and VMs, including high availability, patching, and version upgrades.
Own backup and restore for the services within your area of ownership, using Velero for Kubernetes workloads and the native backup services provided by each cloud platform.
Operate our identity services around Keycloak (SSO, OIDC) and the Kafka streaming platform for event-driven use cases, and help product teams use both safely.
CI/CD, automation & scripting
Build and maintain CI/CD pipelines in Azure DevOps for infrastructure, deployments, and operational tasks within your assigned scope.
Automate provisioning, configuration, and lifecycle operations with Infrastructure as Code (IaC) and a GitOps approach.
Write and maintain scripts and small tools (Python, Bash, PowerShell) that remove manual work from daily operations.
Keep raising the level of automation: when a task comes back regularly, you turn it into a pipeline, a script, or a self-service workflow.
AI infrastructure & AI operations
Deploy and operate selected AI services behind our products and internal AI initiatives, such as Azure AI Foundry, Azure OpenAI endpoints, and vector stores.
Treat AI components within your assigned scope as any other production service: provision them through IaC, deploy them through pipelines, and monitor them properly.
Keep AI operations healthy and affordable by managing capacity, quotas, token consumption, and cost visibility.
Work with our AI and R&D teams to bring new AI use cases safely into production.
Troubleshooting, reliability & incident support
Investigate performance and stability issues across the stack, from Kubernetes and stateful services to pipelines and cloud infrastructure.
Dig into incidents related to the services you support until the root cause is understood, and feed what you learn back into automation, monitoring, and documentation.
Define and track SLIs/SLOs for the services you own and connect them to our central observability stack.
Security, compliance & cost management
Apply TLS/mTLS, RBAC, secrets management, and least-privilege access across all services you operate.
Keep an eye on cost: sensible sizing, storage and retention guardrails, and FinOps practices across environments.
Use of AI to increase efficiency
Use AI-assisted tooling for log analysis, diagnostics, and routine operations work where it genuinely helps.
Involve Security, Compliance, and Data Protection whenever AI touches operational data.
Collaboration & documentation
Work closely with R&D, Cloud Architecture, Security, and the other COS engineers (Database, AI Infrastructure, Observability).
Keep documentation and runbooks current so that knowledge does not live only in your head.
Key Requirements
Several years of hands-on experience operating cloud environments in production, ideally SaaS, on Azure and/or T Cloud Public.
Experience deploying and maintaining a wide range of Azure services, managed as well as self-hosted.
Solid Linux and container experience (Docker, Kubernetes, ideally AKS), including stateful workloads.
CI/CD experience, ideally with Azure DevOps: you have built and maintained pipelines as code, not only used them.
Infrastructure as Code (IaC) and GitOps experience for provisioning and lifecycle management.
Good scripting skills in Python, Bash, or PowerShell, and the habit of automating recurring work.
Hands-on experience with several of: Kafka, Keycloak, Redis, PostgreSQL, Velero.
Strong troubleshooting skills: you work through complex production issues methodically and do not stop at the first symptom.
Solid understanding of networking and security basics (TLS/certificates, RBAC, secrets management).
Experience with or genuine interest in operating AI infrastructure and AI services.
Clear communication with technical and non-technical colleagues; a pragmatic, results-focused mindset.
Fluent English; German is a plus.
Nice-to-have / Plus
Python and Django experience, for example for internal tooling and automation services.
Deeper AI platform experience: model serving, vector databases, AI observability.
Experience with managed counterparts (Azure Event Hubs, Azure Cache for Redis, Azure Database for PostgreSQL) for migration and trade-off decisions.
Deeper identity experience (OIDC, SAML, SSO) with Keycloak.
Experience in sovereign cloud environments such as T Cloud Public.
Relevant certifications (Azure, CKA/CKAD).
Our Benefits
At ATOSS, great talent knows no limits. We welcome professionals from all backgrounds and empower their growth through an inclusive, skill focused environment.
Join us and be part of a high-growth, future-focused company!