← Back to Jobs
Remote
Technology
VS
Staff/Principal-level Linux Networking / HPC Systems Engineer|| Remote
Quick Overview
Seniority
Leader
Work mode
Remote
Location
United States
Posted
6 days ago
TCP/IPBashPython
Job Description
Staff/Principal-level Linux Networking / HPC Systems Engineer
Remote
Phone + Video
Job description:
Remote
Phone + Video
Job description:
Key Responsibilities
Architect, design, and evolve scalable, reliable, and fault-tolerant networking software for high-speed, low-latency interconnects, delivering predictable performance across large-scale MPP systems.
Evaluate and drive adoption of emerging technologies across operating systems, high-performance networking, adapters, DPUs, accelerators, and interconnect fabrics.
Lead complex debugging and root-cause analysis of system-level customer and field issues, including SLES OS crash dump analysis, spanning hardware, firmware, OS, and networking layers.
Define and execute targeted research initiatives and proof-of-concepts to validate new technologies, quantify performance, and guide platform decisions.
Partner with product, hardware, and systems engineering teams to scope, prototype, benchmark, and productionize platform enhancements.
Establish performance benchmarks, validation methodologies, and success metrics for networking and interconnect innovations.
Influence platform roadmaps through deep understanding of industry trends, academic research, and partner technologies.
Mentor and technically guide other engineers through design reviews, code reviews, and architectural discussions.
Leverage AI-assisted coding, analysis, and testing tools to accelerate development cycles and improve code quality and reliability.
Architect, design, and evolve scalable, reliable, and fault-tolerant networking software for high-speed, low-latency interconnects, delivering predictable performance across large-scale MPP systems.
Evaluate and drive adoption of emerging technologies across operating systems, high-performance networking, adapters, DPUs, accelerators, and interconnect fabrics.
Lead complex debugging and root-cause analysis of system-level customer and field issues, including SLES OS crash dump analysis, spanning hardware, firmware, OS, and networking layers.
Define and execute targeted research initiatives and proof-of-concepts to validate new technologies, quantify performance, and guide platform decisions.
Partner with product, hardware, and systems engineering teams to scope, prototype, benchmark, and productionize platform enhancements.
Establish performance benchmarks, validation methodologies, and success metrics for networking and interconnect innovations.
Influence platform roadmaps through deep understanding of industry trends, academic research, and partner technologies.
Mentor and technically guide other engineers through design reviews, code reviews, and architectural discussions.
Leverage AI-assisted coding, analysis, and testing tools to accelerate development cycles and improve code quality and reliability.
Who You’ll Work With
In this role, you will operate across the full lifecycle from research and architecture through production deployment, working closely with platform, hardware, and product engineering teams.
What Makes You a Qualified Candidate
Required Technical Skills:
Strong background in HPC or large-scale distributed systems development.
Proven experience with Linux kernel and driver development in C, including production support.
Deep familiarity with bare-metal and virtualized environments, including performance tradeoffs.
Expertise in InfiniBand and Ethernet networking, leveraging RDMA and RoCE for low-latency, high-throughput communication.
Solid understanding of TCP/IP and UDP networking, along with Linux networking, tuning, and diagnostic tools”.
Packet-level analysis and Linux kernel debugging using tools such as tcpdump, kgdb, and crash.
Experience designing and optimizing high-throughput, low-latency data transport protocols.
Strong knowledge of the Linux kernel, including DKMS, driver lifecycle management, and compatibility across kernel versions.
Proficiency in C, Bash, and Python for systems programming, automation, and diagnostics.
Experience with massively parallel processing (MPP) using message-passing interfaces.
Effective use of modern AI-assisted development tools to accelerate design, coding, and debugging.
In this role, you will operate across the full lifecycle from research and architecture through production deployment, working closely with platform, hardware, and product engineering teams.
What Makes You a Qualified Candidate
Required Technical Skills:
Strong background in HPC or large-scale distributed systems development.
Proven experience with Linux kernel and driver development in C, including production support.
Deep familiarity with bare-metal and virtualized environments, including performance tradeoffs.
Expertise in InfiniBand and Ethernet networking, leveraging RDMA and RoCE for low-latency, high-throughput communication.
Solid understanding of TCP/IP and UDP networking, along with Linux networking, tuning, and diagnostic tools”.
Packet-level analysis and Linux kernel debugging using tools such as tcpdump, kgdb, and crash.
Experience designing and optimizing high-throughput, low-latency data transport protocols.
Strong knowledge of the Linux kernel, including DKMS, driver lifecycle management, and compatibility across kernel versions.
Proficiency in C, Bash, and Python for systems programming, automation, and diagnostics.
Experience with massively parallel processing (MPP) using message-passing interfaces.
Effective use of modern AI-assisted development tools to accelerate design, coding, and debugging.
Nice to Have:
Experience with DPUs, SmartNICs, or hardware offload technologies.
Hands-on work with kernel-bypass networking (e.g., RDMA verbs, DPDK, XDP, eBPF).
Experience with high-speed Ethernet (100G/200G/400G/800G) and modern interconnect fabrics.
Experience tuning systems for NUMA, CPU affinity, cache locality, and memory bandwidth.
Exposure to distributed storage or database platforms in production environments.
Experience working with hardware vendors (NICs, switches, accelerators) on performance or integration issues.
Contributions to open-source networking, kernel, or systems software projects.
Experience with DPUs, SmartNICs, or hardware offload technologies.
Hands-on work with kernel-bypass networking (e.g., RDMA verbs, DPDK, XDP, eBPF).
Experience with high-speed Ethernet (100G/200G/400G/800G) and modern interconnect fabrics.
Experience tuning systems for NUMA, CPU affinity, cache locality, and memory bandwidth.
Exposure to distributed storage or database platforms in production environments.
Experience working with hardware vendors (NICs, switches, accelerators) on performance or integration issues.
Contributions to open-source networking, kernel, or systems software projects.
Education & Experience
Bachelor’s degree in Computer Science (distributed systems focus preferred), Computer Engineering, or Electrical Engineering, or equivalent practical experience.
7+ years of experience in high-performance Linux systems or networking software development, with demonstrated technical leadership.
What You’ll Bring
Confidence and resilience, with the ability to navigate technical disagreement, challenge assumptions, and incorporate feedback constructively.
Proven ability to lead and coordinate real-time troubleshooting of critical (P1) customer issues, rapidly diagnosing system-level failures and driving resolution under pressure.
Strong influencing skills, capable of aligning cross-functional teams and driving outcomes without direct authority or ownership of resources.
Excellent communication skills, with the ability to clearly articulate complex technical findings, remediation plans, and customer impact to both technical and business stakeholders.
A collaborative, self-directed mindset paired with strong intellectual curiosity and continuous learning.
The ability to thrive in ambiguous, fast-paced environments while bringing clarity, structure, and forward momentum.
Bachelor’s degree in Computer Science (distributed systems focus preferred), Computer Engineering, or Electrical Engineering, or equivalent practical experience.
7+ years of experience in high-performance Linux systems or networking software development, with demonstrated technical leadership.
What You’ll Bring
Confidence and resilience, with the ability to navigate technical disagreement, challenge assumptions, and incorporate feedback constructively.
Proven ability to lead and coordinate real-time troubleshooting of critical (P1) customer issues, rapidly diagnosing system-level failures and driving resolution under pressure.
Strong influencing skills, capable of aligning cross-functional teams and driving outcomes without direct authority or ownership of resources.
Excellent communication skills, with the ability to clearly articulate complex technical findings, remediation plans, and customer impact to both technical and business stakeholders.
A collaborative, self-directed mindset paired with strong intellectual curiosity and continuous learning.
The ability to thrive in ambiguous, fast-paced environments while bringing clarity, structure, and forward momentum.
(“ Believe you can and you’re halfway there. ”)
– Theodore Roosevelt
Yogesh Sharma | Lead Tech Recruiter
An -E Verified Company
Similar jobs
- SB
Teamcenter Developer Sunnyvale - CA - California
NewSierra Business Solution LLC
Sunnyvale, CA🇺🇸Hybrid22 hours agoAzureTechnology - OT
AWS Solutions Architect
NewOnwardPath Technology Solutions LLC
Raleigh, NC🇺🇸$65 - $70/hrHybrid22 hours agoDockerMicroservicesAWS+4Technology - SE
Mid-Level Oracle DBA (DBA II)
NewSeersolutionsInc
United States🇺🇸Hybrid22 hours agoOraclePL/SQLSQLTechnology - SE
Senior Oracle DBA (SME)
NewSeersolutionsInc
United States🇺🇸Hybrid22 hours agoOraclePL/SQLSQLTechnology - NS
Workiva Business Data Analyst
NewnTech Solutions
United States🇺🇸Remote22 hours agoTechnology - IS
Business Intelligence Developer: I (Junior)
NewINSPYR Solutions
United States🇺🇸Hybrid22 hours agoSQLETLAgile+1Technology