SDE II - AI Platform
About this role
Most boards and executives are currently flying blind when it comes to cyber risk. They are guessing. At Safe, we’ve built an AI-driven engine that finally gives the C-Suite a clear, quantified, and real-time view of their security posture. We don’t just provide data; we provide certainty.
We are a $170M Series C-funded category leader. We don’t play in the mid-market; we operate at the highest levels of global enterprise. Today, we are proud to serve 10% of the Fortune 500, protecting global icons such as Apple, Netflix, AT&T, Verizon, and Victoria’s Secret.
As we scale toward our next chapter, we are looking for high-performers who want to do the best work of their careers at the intersection of AI and Cybersecurity.
The Culture Memo: Our Operating System
Safe is not a typical corporate environment. We are a high-intensity, mission-driven team. We value builders who want to define a category and work alongside people who are equally committed to excellence.
-
Extreme Ownership: We don’t do "not my job." We hire people who see a gap and own the solution from start to finish.
-
The Elite Standard: We serve the most sophisticated companies on the planet. Our work must be bulletproof. Whether it’s a line of code or a sales deck, we aim for Tier-1 quality every time.
-
Methodology & Rigor: We don’t wing it. From Force Management and MEDDICC in sales to data-driven sprints in engineering, we rely on proven frameworks to stay disciplined and predictable.
-
Radical Candor: We move too fast for politics or sugar-coating. We value direct, honest feedback that helps us find the right answer quickly.
-
The Series C Hustle: We have the stability of a well-funded leader but the heart of a startup.
The Perks & Ownership:
We want our team to feel like owners because they are owners. We trust our people to manage their results and their time.
-
Meaningful Equity: Every "Safestar" is a shareholder. You aren’t just an employee; you are a partner in our success.
-
Unlimited Leaves: We don’t believe in clock-watching. We offer unlimited leave because we trust you to take the time you need to recharge while staying committed to the mission.
-
Comprehensive Benefits: We provide top-tier medical insurance and wellness benefits to ensure you and your family are well cared for.
-
Career Trajectory: We are growing aggressively. For high-performers, the path for advancement moves at the speed of your ambition.
What You'll Do:
-
Build the AI platform: APIs, SDKs, and self-service workflows so AI/ML engineers deploy, version, evaluate, and monitor workloads without touching raw infrastructure.
-
Own GPU infrastructure: Cluster design, provisioning, scheduling, isolation, and utilisation optimisation across shared multi-team demand.
-
Run inference in production: vLLM, Triton, KServe, or Ray — continuous batching, autoscaling, model routing, and multi-model serving against real latency and throughput SLOs.
-
Optimize relentlessly: Drive tokens/sec, p99 latency, GPU utilization, and cost per token through quantization, batching strategy, and capacity planning.
-
Engineer GPU-aware Kubernetes: GPU Operator, device plugins, GPU-aware scheduling, MIG partitioning, and distributed workloads over NCCL.
-
Debug the hard layer: GPU OOMs, driver and CUDA runtime mismatches, interconnect bottlenecks, throughput regressions.
-
Build the MLOps backbone: Model registry and versioning, CI/CD for AI workloads, and safe rollout/rollback.
-
Own the evaluation layer: Offline and online eval harnesses, golden datasets, LLM-as-a-judge scoring, accuracy and hallucination metrics, and automated regression gates so no model, prompt, or quantisation change ships without a measured quality verdict.
-
Make it observable: Response quality and accuracy drift, latency, throughput, GPU utilisation, and per-tenant cost — with SLOs, alerting, and incident response.
-
Secure and multi-tenant by default: AuthN/AuthZ, secrets, data protection, tenant isolation, and governed access to models and GPUs.
-
Codify and lead: Terraform for everything, HA and DR designed in, plus architecture direction and mentoring across teams.
-
Contribute beyond AI: Design and build secure, multi-tenant microservices and APIs on AWS, run thorough code reviews, and own feature delivery end-to-end with Product and Design.
What We're Looking For:
-
2-4 years building and operating production software and infrastructure, with senior ownership of systems end to end.
-
Bachelor's or Master's in Computer Science, Engineering, or equivalent practical experience.
-
Strong Python and/or Go — you write and review production services, not just scripts and manifests.
-
Deep Kubernetes and Docker: scheduling, resource management, operators, networking, debugging under load.
-
Strong Linux, networking, storage, and distributed systems fundamentals.
-
Production AWS/Azure/GCP experience — Lambda, API Gateway, EC2, S3, RDS and equivalents — with Terraform as your default way of working.
-
Working knowledge of SQL and NoSQL databases, including schema design and performance tuning.
-
Experience building and operating backend services and APIs in a multi-tenant SaaS product.
-
Leadership, code review, and mentoring skills, with end-to-end ownership of delivery in an agile environment.
-
Hands-on with a managed ML/AI platform — AWS SageMaker, Bedrock, GCP Vertex AI, or Azure ML — including where it fits and where self-managed infrastructure wins on cost or control.
-
Track record on HA, scalability, and DR in multi-tenant environments.
-
Solid observability and CI/CD practice — metrics, traces, SLOs, automated delivery.
-
Experience building or operating AI evaluation systems — accuracy and quality measurement, LLM-as-a-judge pipelines, eval datasets, and regression testing for models and prompts.
-
Background in platform engineering, infrastructure, SRE, distributed systems, AI infrastructure, or MLOps/LLMOps.
GPU & AI Experience (Required)
-
Hands-on GPU cluster operations: provisioning, capacity planning, and day-2 ownership.
-
NVIDIA ecosystem and CUDA — drivers, container runtime, toolkit compatibility, and their failure modes.
-
GPU scheduling, allocation, isolation, and utilisation optimisation across competing workloads.
-
Kubernetes GPU workloads: GPU Operator, device plugins, GPU-aware scheduling.
-
GPU troubleshooting and tuning: memory/OOM, driver faults, interconnect and throughput bottlenecks.
-
LLM inference and model serving in production (vLLM, Triton, KServe, Ray or equivalent) with demonstrated cost and performance gains.
-
Managed AI platform experience — SageMaker (training jobs, endpoints, inference components) or equivalent — alongside self-managed GPU serving.
Nice to Have:
What Success Looks Like:
-
AI teams deploy and roll back models through self-service workflows, with no manual infrastructure work per deployment.
-
GPU utilisation is measured, forecast, and consistently optimised across the shared fleet.
-
Cost per token falls quarter over quarter, with the numbers on a dashboard.
-
Inference SLOs hold under peak enterprise load, and GPU incidents drop in frequency and time-to-resolve.
-
No model, prompt, or optimisation change reaches production without passing automated accuracy and quality evaluation.
Frequently Asked Questions
Is the salary disclosed for the SDE II - AI Platform position at safe?
Where is the SDE II - AI Platform position at safe located?
Is the SDE II - AI Platform role at safe full-time or part-time?
Which team or department does the SDE II - AI Platform at safe belong to?
How do I apply for the SDE II - AI Platform position at safe?
When was the SDE II - AI Platform job at safe posted?
You'll be redirected to safe's official application page on Lever.