Aitopics
Systems Engineer, Performance
Aitopics, Sunnyvale, California, United States, 94087
Roseland, NJ / Brooklyn, NY / Sunnyvale, CA / Bellevue, WACoreWeave
powers the creation and delivery of intelligence that drives innovation. CoreWeave is the AI Hyperscaler, delivering a cloud platform of cutting-edge services powering the next wave of AI. The company’s technology provides enterprises and leading AI labs with the most performant, efficient, and resilient solutions for accelerated computing. Since 2017, CoreWeave has operated a growing footprint of data centers covering every region of the US and across Europe. CoreWeave was ranked as one of the TIME100 most influential companies of 2024.
As the leader in the industry, we thrive in an environment where adaptability and resilience are key. Our culture offers career-defining opportunities for those who excel amid change and challenge. If you’re someone who thrives in a dynamic environment, enjoys solving complex problems, and is eager to make a significant impact, CoreWeave is the place for you. Join us, and be part of a team solving some of the most exciting challenges in the industry.
About the Role:CoreWeave is seeking a highly skilled and motivated Systems Performance Engineer to join our Kernel
HAVOCK
Team, reporting into the Manager of Systems Engineering. In this role, you will play a crucial part in designing, developing, and optimizing our bare-metal systems from POST through joining a Kubernetes cluster. The team’s primary responsibilities include maintaining a custom Linux kernel, various OS images (Ubuntu-based), the virtualization stack (kubevirt/qemu/vfio), and the container/pod runtime stack (containerd/nydus/kubelet). You will collaborate closely with cross-functional teams, up stack engineering teams, and stakeholders to ensure our low-level software stack is performant in the context of hardware updates; and providing data, metrics, dashboards, and analysis to substantiate performance assertions.
Our Team’s Stack:
Python, Go, bash/sh, C
Intel/AMD/ARM CPUs, Nvidia GPUs, DPUs, Infiniband and Ethernet NICs
Responsibilities:
Develop and maintain tools for establishing systems performance baselines
Develop and maintain performance regression analysis testing automation
Development of telemetry for performance analysis across distributed clusters of servers
Triage and fix performance issues in Linux
Collect data, produce metrics and visualizations that communicate performance information compared to benchmarks; this data should lead to appropriate business decisions and toward greater automation that improves customer experience in relation to performance
Collaborate with cross-functional teams to define Linux and OS requirements, specifications, and system architecture in relation to systems performance
Requirements:
5+ years of professional experience in Systems Performance Engineering
Fluency with a programming language geared toward automation (Python preferred, but others possible)
Experience writing robust, testable code
Experience diagnosing and fixing systems performance issues
Experience with implementing automation testing
Ability to effectively prioritize and communicate proposed features and fixes in a remote-employee environment
Strong passion for automation, with a commitment to automating processes comprehensively
Excellent documentation skills and attention to detail
Strong analytical and problem-solving abilities
Nice-to-haves:
Familiarity with QA/QE best practices
Familiarity with Golang
Opinions about software version control and team collaboration
Experience working in Cloud environments
Experience as a software engineer writing large-scale applications
Experience in open-source community software development
Experience with machine learning is a huge bonus
Compensation:
The base pay for this position ranges from $165,000-$185,000. Pay is based on a number of factors including market location and may vary depending on job-related knowledge, skills, and experience.
Hybrid Workplace:
Successful candidates will be expected to attend onboarding training at our NJ Headquarters within their first several weeks of employment, with subsequent quarterly travel requirements of 1 week duration. If you reside within a 30-mile radius of our New Jersey, New York, or Philadelphia offices, we're excited for you to join us at the office at least three times a week, recognizing the significance we place on fostering connections, collaboration, and creativity within our office culture.
What We Offer:
Medical, dental, and vision insurance - 100% paid for by the employee
Company-paid Life Insurance
Voluntary supplemental life insurance
Short and long-term disability insurance
Flexible Spending Account
Tuition Reimbursement
Mental Wellness Benefits through Spring Health
Family-Forming support provided by Carrot
Flexible, full-service childcare support with Kinside
401(k) with a generous employer match
Flexible PTO
Catered lunch each day in our office and data center locations
A casual work environment
A work culture focused on innovative disruption
Our Workplace:
At CoreWeave, we are committed to operating as a hybrid workplace, offering employees flexibility in how they structure their time between in-office and remote work. We recognize the significance of fostering connections, collaboration, and creativity within our office culture and its positive impact on our business.
CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information.
As part of this commitment and consistent with the Americans with Disabilities Act (ADA), CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship.
#J-18808-Ljbffr
powers the creation and delivery of intelligence that drives innovation. CoreWeave is the AI Hyperscaler, delivering a cloud platform of cutting-edge services powering the next wave of AI. The company’s technology provides enterprises and leading AI labs with the most performant, efficient, and resilient solutions for accelerated computing. Since 2017, CoreWeave has operated a growing footprint of data centers covering every region of the US and across Europe. CoreWeave was ranked as one of the TIME100 most influential companies of 2024.
As the leader in the industry, we thrive in an environment where adaptability and resilience are key. Our culture offers career-defining opportunities for those who excel amid change and challenge. If you’re someone who thrives in a dynamic environment, enjoys solving complex problems, and is eager to make a significant impact, CoreWeave is the place for you. Join us, and be part of a team solving some of the most exciting challenges in the industry.
About the Role:CoreWeave is seeking a highly skilled and motivated Systems Performance Engineer to join our Kernel
HAVOCK
Team, reporting into the Manager of Systems Engineering. In this role, you will play a crucial part in designing, developing, and optimizing our bare-metal systems from POST through joining a Kubernetes cluster. The team’s primary responsibilities include maintaining a custom Linux kernel, various OS images (Ubuntu-based), the virtualization stack (kubevirt/qemu/vfio), and the container/pod runtime stack (containerd/nydus/kubelet). You will collaborate closely with cross-functional teams, up stack engineering teams, and stakeholders to ensure our low-level software stack is performant in the context of hardware updates; and providing data, metrics, dashboards, and analysis to substantiate performance assertions.
Our Team’s Stack:
Python, Go, bash/sh, C
Intel/AMD/ARM CPUs, Nvidia GPUs, DPUs, Infiniband and Ethernet NICs
Responsibilities:
Develop and maintain tools for establishing systems performance baselines
Develop and maintain performance regression analysis testing automation
Development of telemetry for performance analysis across distributed clusters of servers
Triage and fix performance issues in Linux
Collect data, produce metrics and visualizations that communicate performance information compared to benchmarks; this data should lead to appropriate business decisions and toward greater automation that improves customer experience in relation to performance
Collaborate with cross-functional teams to define Linux and OS requirements, specifications, and system architecture in relation to systems performance
Requirements:
5+ years of professional experience in Systems Performance Engineering
Fluency with a programming language geared toward automation (Python preferred, but others possible)
Experience writing robust, testable code
Experience diagnosing and fixing systems performance issues
Experience with implementing automation testing
Ability to effectively prioritize and communicate proposed features and fixes in a remote-employee environment
Strong passion for automation, with a commitment to automating processes comprehensively
Excellent documentation skills and attention to detail
Strong analytical and problem-solving abilities
Nice-to-haves:
Familiarity with QA/QE best practices
Familiarity with Golang
Opinions about software version control and team collaboration
Experience working in Cloud environments
Experience as a software engineer writing large-scale applications
Experience in open-source community software development
Experience with machine learning is a huge bonus
Compensation:
The base pay for this position ranges from $165,000-$185,000. Pay is based on a number of factors including market location and may vary depending on job-related knowledge, skills, and experience.
Hybrid Workplace:
Successful candidates will be expected to attend onboarding training at our NJ Headquarters within their first several weeks of employment, with subsequent quarterly travel requirements of 1 week duration. If you reside within a 30-mile radius of our New Jersey, New York, or Philadelphia offices, we're excited for you to join us at the office at least three times a week, recognizing the significance we place on fostering connections, collaboration, and creativity within our office culture.
What We Offer:
Medical, dental, and vision insurance - 100% paid for by the employee
Company-paid Life Insurance
Voluntary supplemental life insurance
Short and long-term disability insurance
Flexible Spending Account
Tuition Reimbursement
Mental Wellness Benefits through Spring Health
Family-Forming support provided by Carrot
Flexible, full-service childcare support with Kinside
401(k) with a generous employer match
Flexible PTO
Catered lunch each day in our office and data center locations
A casual work environment
A work culture focused on innovative disruption
Our Workplace:
At CoreWeave, we are committed to operating as a hybrid workplace, offering employees flexibility in how they structure their time between in-office and remote work. We recognize the significance of fostering connections, collaboration, and creativity within our office culture and its positive impact on our business.
CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information.
As part of this commitment and consistent with the Americans with Disabilities Act (ADA), CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship.
#J-18808-Ljbffr