Title: Site Reliability Engineer III
Newport News, Virginia, US, 23606
At Jefferson Lab, you’ll champion cutting-edge science and operational excellence while shaping the future of discovery. Join us and make your mark – where excellence meets purpose, and great minds truly matter.
The good-faith pay range for this role is $118,400 - $186,850 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate's skills, experience, education, certifications, and work location.
What your job will be like:
In this job you will:
- Lead the design, implementation, and operation of monitoring, logging, alerting, and diagnostic tooling for HPDF compute, storage, network, and facility systems, contributing directly to that work as well as directing it.
- Supervise, mentor, and develop a small team of site reliability engineers: assign and review work, set expectations, give regular feedback, support technical growth, and plan and estimate the multi-person efforts assigned to the team.
- Establish and maintain the facility's operational framework, including on-call and escalation structure, incident management, change management, and scheduled maintenance, and keep operational records, runbooks, and documentation current.
- Design the facility's resilience model, including failure domain isolation, redundancy, graceful degradation, and disaster recovery objectives for a geographically distributed facility, and validate that design through testing.
- Define, implement, and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) in collaboration with the architecture team and scientific stakeholders, and hold facility operations to them.
- Serve as incident commander for significant incidents, own the postmortem process, and drive root cause prevention back into the design and operation of the systems.
- Drive reliability improvement through automation, process optimization, and the elimination of manual operations, using Python, Go, or shell and standard software development practices.
- Partner with the architecture team on HPDF technology selection from a reliability standpoint, lead evaluations of vendor and open source technologies, and represent HPDF site reliability engineering in the Berkeley Lab partnership.
Additional Responsibilities
- Participate in an on-call rotation as the facility moves toward operations.
Lead - Supervisory - Management
- Supervises a team of site reliability engineers
- Assigns and reviews work, sets performance expectations, provides regular feedback, conducts performance discussions, and supports the technical development of the team.
- Participates in hiring for the group. Does not hold fiscal or budget authority.
Experience
- Required: 10 or more years experience in Site Reliability Engineering, DevOps, systems engineering, or operations engineering, including at least two years leading or supervising engineers. Technical leadership of engineering teams or projects qualifies.
- Preferred: Supporting scientific computing, HPC, or research environments.
- Preferred: Establishing operational practice in a new or greenfield facility.
- Preferred: High availability or around the clock operations.
- Preferred: Serving as the reliability or availability authority during the design phase of a large system or facility, before it entered operations.
- Preferred: Evaluating vendor compute, storage, and network solutions against reliability requirements, including acceptance criteria and benchmarking.
- Preferred: Experience with containers and Kubernetes
- Preferred: Experience with configuration management and infrastructure as code tools (for example Ansible, Terraform, Puppet).
- Preferred: Experience with storage systems, data movement, or large scale data infrastructure.
- Preferred: Experience with IT service management practice and tooling (for example ServiceNow, ITIL).
- Preferred: Experience with HPC infrastructure and environments.
- Preferred: Supporting formal project milestone or gate reviews, such as DOE critical decision reviews, and defining KPPs or acceptance criteria.
Education
- Required: Bachelor's Degree Computer Science or Related Field
- Preferred: Master's Degree Computer Science or Related Field
Experience and Education Exchange
Education above the minimum may be substituted for experience. Relevant experience may not be substituted for education.
Knowledge, Skills, and Abilities
- Deep Linux systems expertise, with the ability to troubleshoot across the application, operating system, storage, and network layers and to guide others in doing so.
- Expertise in designing and operating monitoring and observability stacks (for example Prometheus, Grafana, ELK, OpenTelemetry) and in defining and reporting on SLOs and SLIs.
- Strong scripting and automation skills (Python, Go, or shell) with standard software development practices, together with the judgment to decide what is worth automating.
- Demonstrated ability to lead and mentor technical staff: setting expectations, assigning work, giving constructive feedback, and addressing performance issues in a constructive manner.
- Ability to establish and run incident response and operational process in a production environment.
- Ability to design for resilience, including failure domain isolation, redundancy, graceful degradation, and recovery objectives, and to validate the design through testing and analysis.
- Clear written and verbal communication, including the ability to present options, argue persuasively for proposals, and work productively with a community of scientific users and with colleagues across both laboratories.
- Familiarity with public cloud environments (AWS, Azure, GCP).
- Networking at scale: IPv4/IPv6, DNS, firewalls and access control lists, high speed interconnects, and data transfer protocols.
- Ability to review system and vendor designs from a reliability standpoint and to argue a technical position persuasively with architects, vendors, and scientific stakeholders.
- Load testing, performance analysis, and capacity modeling to validate design assumptions and identify bottlenecks in the data path.
- Ability to estimate cost and effort for multi-person projects and to plan staffing accordingly.
- Practical experience developing and deploying AI assisted or autonomous automation for operational work, and routine use of AI tools in day to day engineering.
About Jefferson Lab
Join a community with a common purpose of solving the most challenging scientific and engineering problems of our time. The Jefferson Lab campus is located in southeastern Virginia amidst a vibrant and growing technology community.
A career at Jefferson Lab is more than a job. You will be part of “big science” and work alongside top scientists and engineers from around the world unlocking the secrets of our visible universe. Managed by SURATech, LLC, Thomas Jefferson National Accelerator Facility is entering an exciting period of mission growth and is seeking new team members ready to apply their skills and passion to have an impact. You could call it work, or you could call it a mission. We call it a challenge. We do things that will change the world.
Total Rewards at Jefferson Lab
At Jefferson Lab, we believe that a comprehensive employee benefits program is an important and meaningful part of the compensation employees receive. Our benefits program includes, but is not limited to:
• Medical, Dental, and Vision Care Plans • Flexible Spending Accounts
• Paid Time-off and Leave Programs (Paid Parental, vacation, holidays, and sick leave)
• 401(k) Plan – 9% Lab Contribution; 100% vested • Flexible Work Arrangements
(Remote & Alternate Work Schedules available)
• Tuition Assistance, Training and Professional Development Programs
• Live near the waterways of the Chesapeake Bay region with access to nearby beaches,
mountains, and all major metropolitan centers on the East Coast
SURATech, LLC manages and operates the Thomas Jefferson National Accelerator Facility (Jefferson Lab). SURATech is an Equal Opportunity Employer.
SURATech is committed to providing reasonable accommodation for people with disabilities (unless doing so will result in an undue hardship). If you need a reasonable accommodation for any part of the employment process, please send an e-mail to recruiting@jlab.org or contact Human Resources by calling (757) 269-7100 and selecting option 1 between 8 am – 5 pm EST to provide the nature of your request.
Employment with SURATech is conditional upon DOE approval if at any time during your employment you are participating in a Foreign Government Talent Recruitment Program or Affiliated activity. Generally, such programs/activities include any foreign-state-sponsored attempt to acquire U.S.-funded scientific research through programs run or funded by the government that target scientists, engineers, students, academics, researchers, and entrepreneurs of all nationalities working or educated in the United States. This includes positions or appointments, both domestic and foreign, titled academic, professional, or institutional appointments whether or not remuneration is received and whether full-time, part-time or voluntary.
Nearest Major Market: Hampton Roads