Platform Core: Design and operate large-scale distributed data systems
Own the big data compute and storage infrastructure (MaxCompute/ODPS, Hologres, Spark)
Build and maintain multi-site task orchestration that dynamically selects engines and enforces policy
Drive reliability and performance improvements across batch and real-time pipelines
AI Integration: Build the AI-native platform layer
Develop and expose MCP (Model Context Protocol) tool interfaces so AI agents can interact with platform APIs
Build the scheduling and cost-optimization agents that auto-tune resource allocation and alert severity
Instrument platform telemetry to feed AI-driven SLA monitoring and anomaly detection
Design context retrieval pipelines (RAG / vector search) for SQL code and config knowledge bases
Tooling & DX: Evolve the developer experience
Own the internal data development platform — IDE integrations, code review automation, deployment tooling
Build APIs-first tools (backfill, ingestion automation) designed for future MCP integration
Collaborate with data warehouse and service teams to define platform contracts
Ops & Governance: Drive operational excellence
Establish SLA benchmarks, cost metrics, and latency dashboards as AI optimization targets
Build automated incident response and root-cause analysis pipelines
Define and enforce infrastructure policies across multi-cloud environments
AI Agent Ownership — Platform Tier
Scheduling Agent: auto-configure task dependencies, engine selection, cost/performance trade-offs, and alert tiers
Operations Agent: detect pipeline latency, performance degradation, and schema drift; trigger remediation
Incident Response Agent: trace SLA breaches to root cause, assign accountability, generate post-mortems
MCP Tool Layer: design and maintain the cross-platform tool interfaces that all agents call into