Design and own database architecture for critical infrastructure and platform services, including PostgreSQL-backed internal platforms, Slurm accounting and operational databases, NetBox and infrastructure source-of-truth databases, custom internal applications and automation services, observability, inventory, and platform metadata systems, future database-backed control plane services
Define standard database patterns for high availability, replication, failover, backup and restore, point-in-time recovery, performance baselining, capacity planning, upgrade lifecycle management, access control and operational security
Establish database design standards for new internal platforms, including schema review, indexing strategy, query design, service ownership boundaries, and production readiness requirements
Operate and improve production database environments across PostgreSQL, MySQL, Percona, and adjacent systems
Own the lifecycle of database systems, including provisioning, configuration, version upgrades, replication topology design, performance tuning, backup validation, disaster recovery testing, decommissioning, documentation and runbook creation
Troubleshoot and resolve production database issues involving query latency, lock contention, replication lag, storage I/O bottlenecks, connection exhaustion, poor indexing, schema design problems, database capacity constraints, backup or restore failures
Drive root cause analysis for database-related incidents and convert findings into durable engineering improvements
Build deep database observability beyond basic dashboards
Develop and maintain visibility into query performance, execution plans, index usage, replication health, locking behavior, buffer/cache efficiency, storage latency, connection pool behavior, OS-level database bottlenecks
Use tools such as PostgreSQL native statistics, MySQL/Percona tooling, Prometheus, Grafana, PMM, Query logs, slow query logs, eBPF/BCC or equivalent low-level profiling tools, Linux performance tooling
Create performance baselines and alerting standards for critical database platforms
Identify recurring database failure patterns and build preventive monitoring, automation, and operational guardrails
Create database automation patterns that can be integrated with existing infrastructure tooling
Partner with DevOps and Infrastructure Engineering to automate database provisioning, configuration standards, backup verification, health checks, replication checks, user and permission management, upgrade workflows, monitoring deployment, runbook-driven recovery procedures
Contribute database-specific modules, roles, or workflows to Ansible, CI/CD pipelines, or internal automation platforms where appropriate
Define production database readiness standards for new services before they are promoted into critical environments