The Systems Performance team is part of the Computing Squad consists of two distinct workstreams —Orchestration and System Performance —each with its own approach and challenges on managing and improving the foundational infrastructure where the majority of the Nubank's workloads runs
The Performance team is focused on building deep diagnostic tools and performing high-level analysis to reduce latency, infrastructure costs and increase services efficiency
You will be responsible for leading complex performance investigations, identifying systemic bottlenecks, and driving efficiency across one of the largest JVM-based microservice architectures in the world
Our core principles and behaviors include ownership, simplicity, veracity-first, teamwork, and a focus on quality over quantity. During a normal work day, you will interact with critical infrastructure layers, from the Linux Kernel and JVM internals to cloud-wide orchestration
Leading Deep-Dive Investigations: Conduct high-level performance analysis to identify and resolve systemic bottlenecks across our global JVM-based microservices architecture
Optimizing Resource Efficiency: Drive initiatives to reduce infrastructure costs and latency by fine-tuning JVM parameters, Garbage Collection (ZGC, G1), and memory management (heap and off-heap)
Building Diagnostic Tooling: Develop and implement advanced observability tools using eBPF, JFR, and Flamegraphs to provide real-time insights into kernel and runtime behavior
Kernel & Runtime Alignment: Bridge the gap between the Linux Kernel and the JVM, optimizing thread scheduling (CFS/EEVDF) and managing resource isolation (cgroups/throttling) within our Kubernetes environment
Architecting Scalable Solutions: Design and deliver innovative infrastructure improvements that address long-term performance challenges, ensuring our systems scale ahead of demand
Technical Mentorship & Culture: Share expertise on JVM internals and performance best practices with the wider Engineering team, fostering a culture of technical excellence and "quality over quantity."
Root Cause Excellence: Deep dive into complex concurrency issues, lock contentions, and memory leaks, providing definitive fixes for high-impact technical debt
Strategic Collaboration: Work closely with the Computing Squad to align orchestration strategies with system performance goals, ensuring a seamless interface between infrastructure and workloads