RV LIFE is looking for a Senior DevOps & Infrastructure Lead to help us stabilize, document, and modernize the infrastructure behind our products
This is a hands-on senior role for someone comfortable inheriting real production systems, reducing operational risk, improving reliability, and moving us toward a documented, secure, automated, infrastructure-as-code operating model
We run production across DigitalOcean, AWS, Cloudflare, and other hosting providers, and are consolidating onto managed, infrastructure-as-code platforms. We need deep, hands-on expertise across these environments
RV LIFE is an AI-first engineering organization. We expect this person to use AI to accelerate discovery, documentation, runbooks, log review, scripting, and infrastructure-as-code drafting, while applying strict human judgment around security, secrets, production access, destructive commands, rollback, and correctness
This role focuses on the infrastructure path to reliability; application-level architecture changes are handled in partnership with our engineering team. It is not just about keeping servers alive. It is about building durable practices that reduce single-person dependency, improve visibility, and make our systems safer to operate
This is not a standard 9-to-5 role. Production issues do not keep business hours, so it carries real on-call responsibility: you need to be reachable and able to respond when unforeseen incidents arise
Administer and improve existing DigitalOcean infrastructure
Support and improve Linux-based production server environments
Migrate self-managed databases onto managed database services, with validated failover, backups, and recovery
Move applications onto managed runtimes (including Laravel Cloud where it fits), replacing manual deploy processes with automated, repeatable pipelines
Expand and harden our use of Cloudflare for edge, static hosting, caching, and security
Build a clear inventory of servers, services, databases, domains, access paths, backups, monitoring, and operational risks
Create and maintain practical runbooks for common and emergency infrastructure workflows
Improve incident response, escalation paths, monitoring, logging, and alerting
Review and improve backup, restore, and disaster-recovery procedures
Identify recurring manual work and convert it into safer procedures, scripts, automation, or infrastructure-as-code
Help define infrastructure-as-code standards and move appropriate infrastructure into repeatable, version-controlled workflows
Work with AWS services where needed (Lambda, VPC, IAM, CloudWatch, S3, SSM/Secrets Manager, queues)
Use AI tools to accelerate discovery, documentation, scripting, troubleshooting, and automation, with strong production-safety judgment
Partner with engineering leadership to prioritize infrastructure risk and modernization; track work clearly in Jira/GitHub and communicate proactively about risks, tradeoffs, and blockers
What Success Looks Like
In the first 30-60 days, you'll take ownership of how we see and operate our infrastructure, building on what we already track and closing the gaps
You'll validate and take ownership of what already exists
Our infrastructure inventory and server map
Our monitoring and alerting
Our DNS / Cloudflare configuration
Our prioritized infrastructure risk register