NOC Technician (Data Center and Site Ops)

xai· Data Center
Apply Now ↗
📍 Memphis, Tennessee

About this role

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

As a NOC Technician, you are the eyes and the voice of the site — never the hands. You staff the Network / Campus Operations Center and continuously observe site health signals across xAI campuses. You detect and verify campus-impacting events, assemble the right responders, run incident communications leadership can trust, and drive every major incident to a completed report and a tracked corrective project. You work with Site Reliability Engineering, SiteOps, Facilities, Hardware Failure Analysis, SWE Platforms, and vendors — escalating correctly the first time and maintaining the institutional memory across shifts and sites.
One sentence: watch the campus, run the bridge, leave the wrench work and deep root cause to the teams that own them.

RESPONSIBILITIES:
Continuous monitoring (the watch)
• Staff the console per shift schedule to sustain 24/7 coverage (coverage posture: 2 on console per site)
• Watch the designated signal surface: cluster health dashboards, node availability, network health, facility trend panels (power/cooling), storage alarms, and threshold breaches as defined by SRE monitoring standards
• Acknowledge every page/alert within the SLA; classify it (actionable / known / noise) and log the disposition; feed noise patterns back to SRE for suppression or redesign
• Maintain a live picture of ongoing maintenance, planned work, and degraded-but-accepted states so real anomalies stand out

Detection, triage & escalation
• Detect → verify → escalate within defined time budgets; verification is signal-level (is it real, what's the blast radius), not deep diagnosis
• Operate the escalation matrix: NOC → on-call SRE → domain owners (SiteOps, Facilities, Network, Storage, HW FA, vendors); page correctly the first time
• Recommend incident declaration and severity to the on-call SRE; declare directly per runbook when thresholds are unambiguous

Incident communications & coordination
• Open and run the bridge; get the right people on within the time-to-bridge SLA
• Own stakeholder communications: first update within the SLA, then a fixed cadence until resolution
• Maintain the incident timeline in real time — timestamps, actions, decisions, engagements
• Track who owns what during the incident and call out stalls

First-pass RCA framing & closure
• Produce initial framing for major site outages: what happened, when it started, what's impacted (halls/racks/services), what changed recently, who is engaged
• Hand framing to SRE / Hardware FA for depth — the NOC does not publish root cause
• Write major-incident reports; open corrective projects in Linear with named owners and track them to closure ("filed" is not "done")

Shift operations, runbooks & improvement
• Run structured shift handoffs and keep durable shift logs; maintain cross-site awareness
• Own and continuously improve NOC runbooks: escalation matrix, comms templates, severity ladders, per-signal response procedures
• Participate in game days run by SRE; every incident where the runbook was wrong or missing produces a runbook change before the incident closes

Explicitly not this role
• Wrench work: swaps, reseats, physical recovery (SiteOps)
• Power / cooling / building plant operation (Facilities)
• Deep hardware root-cause analysis or vendor CAPA (Hardware Failure Analysis)
• Monitoring architecture, alert design, or technical SEV command (Site SRE)
• Building or operating reliability tooling such as SRT, turnback, or dashboards (SWE Platforms)

BASIC QUALIFICATIONS:
• High school diploma or equivalency certificate
• 1+ year of professional experience in a Network Operations Center (NOC), Security Operations Center (SOC), mission-control / dispatch, data center operations watch, or equivalent 24/7 monitoring and incident-communications role
• Demonstrated written and verbal communication skills under time pressure (stakeholder updates, handoffs, timelines)

PREFERRED SKILLS AND EXPERIENCE:
• Calm under pressure; excellent written and verbal communications — leadership should be able to trust your incident updates verbatim
• Pattern recognition across domains; multi-domain curiosity (compute, network, storage, power/cooling signals)
• Experience following and improving process: runbooks, escalation matrices, shift handoffs, post-incident follow-through
• Prior NOC, SOC, or critical-environment operations experience in a datacenter or hyperscale infrastructure environment
• Familiarity with reading operational dashboards, acknowledging/classifying alerts, and coordinating across on-site technicians, facilities, and engineering on-call
• Comfort with ticketing / project tracking systems (e.g. Linear, Jira) for opening and chasing corrective work to closure
• Industry certifications a plus (Network+, Security+, ITIL, or similar) — not a substitute for judgment and communications quality
• Basic familiarity with datacenter topology (racks, fabric, OOB) and how facility events affect compute availability — enough to triage and escalate correctly, not to deep-diagnose
• Bachelor's degree in IT, Computer Science, Cybersecurity, or STEM discipline preferred but not required

ADDITIONAL REQUIREMENTS:
• Must be available for on-shift rotations supporting 24/7/365 console coverage
• Shift structure (e.g. 12-hour rotations) to be confirmed; nights, weekends, and holidays are part of the role
• Must be able to work extended hours during major incidents as needed

SUCCESS LOOKS LIKE:
• Coverage attainment / shift fill rate vs plan
• Time-to-bridge for major incidents; first-update and cadence SLA attainment
• Page accuracy / escalation correctness
• % of major incidents with a complete timeline and follow-up projects tracked to done
• Not measured by: raw page counts, or heroics without a paired prevention item

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Frequently Asked Questions

Is the salary disclosed for the NOC Technician (Data Center and Site Ops) position at xai?
The salary for this NOC Technician (Data Center and Site Ops) role at xai is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the NOC Technician (Data Center and Site Ops) position at xai located?
This NOC Technician (Data Center and Site Ops) role at xai is based in Memphis, Tennessee. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Which team or department does the NOC Technician (Data Center and Site Ops) at xai belong to?
This NOC Technician (Data Center and Site Ops) position is part of the Data Center department at xai. See the full job description for more information about the team structure and responsibilities.
How do I apply for the NOC Technician (Data Center and Site Ops) position at xai?
Click the "Apply Now" button on this page. You will be redirected to xai's official application portal hosted on greenhouse where you can submit your application directly.
When was the NOC Technician (Data Center and Site Ops) job at xai posted?
This NOC Technician (Data Center and Site Ops) position at xai was posted on Aug 13, 2026. Apply as soon as possible — early applications are often reviewed first.
NOC Technician (Data Center and Site Ops)
xai
Apply for this role ↗

You'll be redirected to xai's official application page on Greenhouse.