Designing and owning the end-to-end MLOps architecture for production Machine Learning and Generative AI systems
Building and maintaining ML training, validation, deployment, serving, monitoring, and retraining pipelines
Defining and implementing the model promotion lifecycle across development, staging, and production environments
Building and maintaining model registry, experiment tracking, dataset/model versioning, and reproducible ML workflows
Designing scalable training and inference infrastructure, including GPU-backed workloads
Building and maintaining CI/CD pipelines for ML and AI workloads with automated testing, quality gates, approval controls, and rollback mechanisms
Implementing progressive delivery approaches for ML models, including shadow testing, canary releases, and blue-green deployments
Collaborating closely with Platform Engineering to run AI workloads on shared Kubernetes infrastructure
Translating ML infrastructure requirements into technical requirements for platform, security, capacity, and architecture teams
Establishing reusable MLOps project templates, shared pipeline components, and engineering standards
Implementing automated ML testing, including data validation, model regression tests, training-serving consistency checks, and evaluation gates
Owning production reliability of ML systems, including availability, latency, throughput, scalability, and operational stability
Building observability across infrastructure, data quality, model performance, and business impact
Configuring monitoring, alerting, incident response, runbooks, rollback procedures, and post-incident reviews for AI systems
Implementing automated model retraining based on schedules, events, and data/model drift
Managing the full model lifecycle, including deployment, monitoring, retraining, version promotion, and retirement
Tracking and optimizing infrastructure, training, inference, and GPU-related costs
Building and operating the production layer for Generative AI and LLM-based applications
Implementing LLM model gateways, routing, prompt management, prompt versioning, caching, and rate limiting
Implementing LLM guardrails, groundedness monitoring, sensitive-data protection, and prompt-injection mitigation
Building observability and permission controls for agentic AI systems
Implementing token-level and request-level consumption metering across Generative AI applications
Building cost attribution mechanisms by business unit, use case, application, model, tenant, and environment
Optimizing LLM workloads through model routing, prompt/context optimization, caching, batching, and selection of cost-efficient models
Applying hands-on Generative AI engineering practices, including prompt engineering, structured outputs, context management, embeddings, tool/function calling, and model adaptation where required
Evaluating emerging ML and Generative AI tools, models, and platforms and recommend technologies for adoption
Making and documenting architectural and build-vs-configure decisions for new AI capabilities
Producing architecture decision records, reference designs, technical documentation, and operational guidelines
Participating in architecture, capacity planning, technical roadmap, and platform strategy discussions
Performing code reviews, technical mentoring, pairing, and knowledge sharing to improve engineering standards across the team
Strong engineering experience across Software Engineering, DevOps/SRE, and MLOps for 5+ years
Production ownership of machine learning systems: you have operated, supported, and been accountable for ML in production
Deep hands-on experience with Docker and Kubernetes, including resource management and GPU workloads
Strong proficiency with GitLab CI, GitHub Actions, ArgoCD, or equivalent, including pipeline-enforced quality gates
Extensive experience with MLflow, Kubeflow, Feast, BentoML, KServe, SageMaker, Vertex AI, Databricks, or comparable platforms
Python at production engineering standard (typed, tested, packaged, and reviewed)
Infrastructure as Code (Terraform, Ansible, or equivalent, with real module and state management experience)
Production service engineering (stateless, configuration-driven services with structured logging, published API contracts, and centrally managed secrets)
Data services in production (PostgreSQL, caching, and message-driven asynchronous Processing)
Monitoring and observability (metrics, logging, error tracking, and ML-aware monitoring such as Evidently, WhyLabs, or Arize)
Proven collaboration with platform or infrastructure engineering teams, building on shared infrastructure rather than around it
Hands-on LLMOps experience: model gateways, prompt versioning, RAG pipelines, vector databases, evaluation harnesses, and guardrail frameworks
Generative AI consumption management: token metering, cost attribution across multiple applications or tenants, budget and quota controls, and demonstrable cost optimization of LLM workloads
Model evaluation and selection: building evaluation suites that objectively compare models on quality, latency, cost, and safety for a given task, and using them to drive adoption decisions
Practical generative AI engineering: prompt and structured output design, context management, embedding selection, tool calling, and fine-tuning or model adaptation
Experience operating agentic AI systems in production
Experience establishing an ML platform capability from an early stage
Cost engineering for AI workloads: GPU efficiency, inference optimization, and consumption attribution
Familiarity with AI governance and regulated environments — model risk management, auditability, and fairness testing
Experience with data quality frameworks and workflow orchestration
Level of English – Upper-Intermediate and above
Experience in teamwork with leaders in FinTech, Healthcare, Retail, Telecom, and others. Andersen cooperates with such businesses as Samsung, Siemens, Johnson & Johnson, BNP Paribas, Ryanair, Mercedes, TUI, Verivox, Allianz, T-Systems, etc
The opportunity to change the project and/or develop expertise in an interesting business domain
Job conditions – you can work both fully remotely and from the office or can choose a hybrid variant
Guarantee of professional, financial, and career growth! The company has introduced systems of mentoring and adaptation for each new employee
The opportunity to earn up to an additional 1,000 EUR per month, depending on the level of expertise, which will be included in the annual bonus, by participating in the company's activities
Access to the corporate training portal, where the entire knowledge base of the company is collected and which is constantly updated
Bright corporate life (parties / pizza days / PlayStation / fruits / coffee / snacks / movies)
Certification compensation (AWS, PMP, etc)
Referral program
Private health insurance and sports compensation, depending on the type of employment