Cloud and AI Engineer
CHAPTERS Group AGCologneArbeitnow١٦/٩/٢٠٢٦
إعلان
About CHAPTERS Group
CHAPTERS Group AG is a Hamburg-based holding company investing in 55+ mission-critical software companies across the DACH region, France, and the Czech Republic. Our companies deliver the digital solutions that keep the public sector and core industries running. What they do matters in daily life.
AI has rapidly disrupted the software industry, and we aim to lead that transformation across our portfolio. Central to this is the AI Core team: we build and operate a shared enterprise AI Hub used by companies across the group, and drive adoption through a network of AI Champions. Each portfolio company has one designated power user who leads AI enablement locally, runs training, and scales usage within their organisation.
Tasks
The Role
You are a core member of the AI Core team, owning the infrastructure that powers AI adoption across 55+ portfolio companies. Your focus is the AI Hub platform (Open WebUI): you keep it running, deploy new instances, build and validate integrations at HQ level, and make sure every company in the group gets a stable, well-configured environment.
We build the AI Hub as a product and support to ship it in every company. Each company owns their HUB and administers it (L1 Support). You are part of the central team that drives features and changes that are of high value for all hubs in the portfolio.
In addition, you are part of the team that serves as the center of expertise (L2 Support) and give guidance and hands-on support when the first level teams need you.
What You Will Work On
Platform engineering: new integrations, container images, infrastructure-as-code, deployment tooling and security patches. This is the product that every company runs
Release qualification: when the upstream project ships a new version, you assess it (one-way database migrations, breaking changes, patch compatibility), validate it on a canary, and publish a qualified release with notes that an L1 can act on
Shared infrastructure: Application Gateway, WAF policies, TLS certificates, Redis, DNS and network peering. This layer is shared across every company and cannot be devolved. A mistake here affects all of them
Enable the L1 network and be their technical point of contact: write the developer guides, run onboarding and office hours, and build preflight checks and guardrails that make it safe for a platform admin to deploy and operate their own Hub.
Integrations Catalog stewardship: versioned artefacts (tools, skills, agents, integrations) each shipped with a user guide for AI Champions and a developer guide for platform admins
Fleet observability: reporting and alerting version parity, uptime, cloud spend and deployment status across all companies.
Own the eval service: the service itself, the quality of judges and generators, and the methodology platforms use to measure whether their agents do the work
Close the loop. Shipping is the beginning. Your designs gather operational signals from the fleet and feed a continuous improvement cycle
Requirements
What We Are Looking For
Required
At least 2 years of professional experience in platform engineering, DevOps, SRE, or cloud operations, including ownership of production systems that others depend on
Production cloud experience on Azure or AWS
Infrastructure as Code in production: Bicep, Terraform, CDK or CloudFormation. We care that you think in IaC.
Strong Python skills beyond scripting: you have built and maintained small production services, worked with APIs, and found your way around unfamiliar codebases
Comfortable with Linux, Docker, Git and CI/CD as everyday working tools
Working knowledge of databases and identity: PostgreSQL in production, and OAuth / SSO / Entra ID concepts
AI-native: you use LLMs as tools, understand their operational properties, and have shipped something with them
English (fluent): primary working language across the team; German conversational is a plus
Bonus
Clear technical writing: Enabling others is a real pa
