Logo
Aitopics

Systems Engineer, Performance

Aitopics, Sunnyvale, California, United States, 94087


Roseland, NJ / Brooklyn, NY / Sunnyvale, CA / Bellevue, WACoreWeave

powers the creation and delivery of intelligence that drives innovation. CoreWeave is the AI Hyperscaler, delivering a cloud platform of cutting-edge services powering the next wave of AI. The company’s technology provides enterprises and leading AI labs with the most performant, efficient, and resilient solutions for accelerated computing. Since 2017, CoreWeave has operated a growing footprint of data centers covering every region of the US and across Europe. CoreWeave was ranked as one of the TIME100 most influential companies of 2024.

As the leader in the industry, we thrive in an environment where adaptability and resilience are key. Our culture offers career-defining opportunities for those who excel amid change and challenge. If you’re someone who thrives in a dynamic environment, enjoys solving complex problems, and is eager to make a significant impact, CoreWeave is the place for you. Join us, and be part of a team solving some of the most exciting challenges in the industry.

About the Role:CoreWeave is seeking a highly skilled and motivated Systems Performance Engineer to join our Kernel

HAVOCK

Team, reporting into the Manager of Systems Engineering. In this role, you will play a crucial part in designing, developing, and optimizing our bare-metal systems from POST through joining a Kubernetes cluster. The team’s primary responsibilities include maintaining a custom Linux kernel, various OS images (Ubuntu-based), the virtualization stack (kubevirt/qemu/vfio), and the container/pod runtime stack (containerd/nydus/kubelet). You will collaborate closely with cross-functional teams, up stack engineering teams, and stakeholders to ensure our low-level software stack is performant in the context of hardware updates; and providing data, metrics, dashboards, and analysis to substantiate performance assertions.

Our Team’s Stack:

Python, Go, bash/sh, C

Intel/AMD/ARM CPUs, Nvidia GPUs, DPUs, Infiniband and Ethernet NICs

Responsibilities:

Develop and maintain tools for establishing systems performance baselines

Develop and maintain performance regression analysis testing automation

Development of telemetry for performance analysis across distributed clusters of servers

Triage and fix performance issues in Linux

Collect data, produce metrics and visualizations that communicate performance information compared to benchmarks; this data should lead to appropriate business decisions and toward greater automation that improves customer experience in relation to performance

Collaborate with cross-functional teams to define Linux and OS requirements, specifications, and system architecture in relation to systems performance

Requirements:

5+ years of professional experience in Systems Performance Engineering

Fluency with a programming language geared toward automation (Python preferred, but others possible)

Experience writing robust, testable code

Experience diagnosing and fixing systems performance issues

Experience with implementing automation testing

Ability to effectively prioritize and communicate proposed features and fixes in a remote-employee environment

Strong passion for automation, with a commitment to automating processes comprehensively

Excellent documentation skills and attention to detail

Strong analytical and problem-solving abilities

Nice-to-haves:

Familiarity with QA/QE best practices

Familiarity with Golang

Opinions about software version control and team collaboration

Experience working in Cloud environments

Experience as a software engineer writing large-scale applications

Experience in open-source community software development

Experience with machine learning is a huge bonus

Compensation:

The base pay for this position ranges from $165,000-$185,000. Pay is based on a number of factors including market location and may vary depending on job-related knowledge, skills, and experience.

Hybrid Workplace:

Successful candidates will be expected to attend onboarding training at our NJ Headquarters within their first several weeks of employment, with subsequent quarterly travel requirements of 1 week duration. If you reside within a 30-mile radius of our New Jersey, New York, or Philadelphia offices, we're excited for you to join us at the office at least three times a week, recognizing the significance we place on fostering connections, collaboration, and creativity within our office culture.

What We Offer:

Medical, dental, and vision insurance - 100% paid for by the employee

Company-paid Life Insurance

Voluntary supplemental life insurance

Short and long-term disability insurance

Flexible Spending Account

Tuition Reimbursement

Mental Wellness Benefits through Spring Health

Family-Forming support provided by Carrot

Flexible, full-service childcare support with Kinside

401(k) with a generous employer match

Flexible PTO

Catered lunch each day in our office and data center locations

A casual work environment

A work culture focused on innovative disruption

Our Workplace:

At CoreWeave, we are committed to operating as a hybrid workplace, offering employees flexibility in how they structure their time between in-office and remote work. We recognize the significance of fostering connections, collaboration, and creativity within our office culture and its positive impact on our business.

CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information.

As part of this commitment and consistent with the Americans with Disabilities Act (ADA), CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship.

#J-18808-Ljbffr