Lawrence Berkeley National Laboratory

Network Engineer, Platform, Automation & HPC/AI

San Francisco Bay Area, CA, US$139,440-$267,996Posted 13 days ago

Job Description

Join NERSC at Berkeley Lab and help engineer the high-performance network platform powering some of the nation’s most advanced scientific computing. As a Network Engineer, Platform, Automation & HPC/AI, you’ll work at the intersection of network engineering, automation, and software development while helping advance AI-driven network operations. Our team manages 1 Tb/s of border connectivity to ESnet and an 800G/400G data center network backbone supporting the NERSC-9 and NERSC-10 supercomputers, multi-tier storage, archive, and edge services. Your work will help improve the performance, scalability, automation, and reliability of scientific workflows serving more than 10,000 users.

Our mission is to bring science solutions to the world. We welcome candidates from all backgrounds, including those with non-traditional paths. We value a growth mindset and believe skills and experience are transferable. If you’re eager to learn and meet the qualifications below, we encourage you to apply. Join our team where your work can have a high impact for an organization associated with 17 Nobel Prizes… and counting.

This position may be filled at Level 3 or Level 4. Level 3 is intended for experienced engineers who independently solve complex networking and automation challenges. Level 4 is intended for senior technical leaders who architect solutions, lead modernization efforts, and establish new technical approaches for complex HPC and data center environments.

Network Engineer Level 3 will

* Implement, operate, maintain, and improve network automation and observability solutions. * Contribute to Data Center modernization efforts and NERSC’s Smart Facility initiative. * Support efforts to design and deliver network services to address emerging needs (e.g., American Science Cloud, new Edge services). * Continuously monitor and optimize network performance, focusing on reducing latency, maximizing throughput, and improving fault tolerance. * Create and

Apply for this role

Keep looking

Similar Remote AI Jobs