Обязанности
•Build CI CD pipelines and container deployments
•Build observability dashboards and alerts
•Build prompt agent and model release pipelines
•Conduct load resilience and red team testing
•Conduct threat modeling and security reviews
•Define service level objectives and escalation processes
•Develop incident playbooks and runbooks
•Develop unit economics for AI services
•Embed DevSecOps practices and secure development lifecycle
•Ensure secure handling of secrets and sensitive data
•Implement automated security testing in CI/CD
•Implement automated testing and evaluation pipelines
•Implement distributed tracing metrics logging with OpenTelemetry
•Implement model routing and fallback
•Investigate and respond to AI service incidents
•Maintain dependency vulnerability and supply chain controls
•Maintain incident records and post-incident reviews
•Maintain model and provider version controls
•Manage AI provider integrations
•Manage canary deployments and rollbacks
•Manage secrets authentication networking and access controls
•Monitor model drift and output quality
•Operate model gateway
•Optimize cost with routing caching batching and context reduction
•Set rate limits and quotas
•Support auditability and data protection compliance
•Use Infrastructure as Code for AI infrastructure
Условия
•Career development support
•Cross-functional collaboration
•Flexible ownership of AI platform
•Industry certification support
Технологии: AI Services, Alerting, Application Insights, Audit Logging, Authentication, Automated security, Automated security testing, Azure AI, Azure AI Services, Azure Application Insights, Azure DevOps, Azure Key Vault, Azure OpenAI, Batching, CI/CD, Caching, Canary deployments, Cloud infrastructure, Containers, Dashboards, DevSecOps, Distributed tracing, Encryption, Fallback strategies, Incident Response, Infrastructure as Code, Key Vault, Latency optimization, Logging, Metrics, Model Evaluation, Model routing, Networking, Observability, OpenTelemetry, Prompt engineering, Python, Quotas, Rate Limiting, Red team, Red team testing, Rollback procedures, Secrets management, Security Testing, Serverless, Service Level, Service-Level Objectives, Software Supply Chain, Software supply chain security, Supply chain security, Threat modeling, Token Usage Monitoring, Token usage, Usage monitoring, Vulnerability Management, “as-code”