San Francisco, CA · On-site · Full-time
Compensation: $150,000–$250,000 + up to 1% equity
A seed-stage LLM interpretability and context-optimization company building custom machine learning models that analyze and compress token contexts before they reach the underlying model — delivering roughly 50% inference cost reduction, lower latency, and higher accuracy for the enterprises and scale-ups integrating LLMs into their products. Venture-backed, with strong early traction (~1,000 customers within its first seven months).
Founded 2025 · 1–10 people · Industry: AI Tools / LLM infrastructure
Own the full multi-region GPU infrastructure stack end to end as the sole infra hire — global low-latency serving, multi-cloud and on-premise deployments, reliability, and cost efficiency. Your work sits directly in the critical path of live customer traffic. This is a high-ownership, in-person role at a fast-moving early-stage team, working a 996 pace (9am–9pm, six days a week) in San Francisco.
What you'll be doing
-
Own the full infrastructure stack end to end across multi-region GPU deployments, major public clouds, and on-premise enterprise environments.
-
Build and maintain super low-latency GPU serving infrastructure that sits in the critical path of live customer traffic.
-
Manage multi-cloud deployments, including cloud marketplace integrations and provider relationships.
-
Design and iterate on deployment, scaling, reliability, and cost-efficiency systems as the sole infra owner.
-
Support on-premise deployments for enterprise clients and ensure performance and reliability at each site.
-
Research and adopt new infrastructure solutions continuously as the stack and customer base grow.
Tech stack: AWS, GCP, Terraform, Docker, CI/CD, GPU/ML inference infrastructure
-
Own cloud systems serving compression API end-to-end
-
Build and operate global low-latency high-throughput GPU ML inference infrastructure
-
Work with AWS, Terraform, Docker and CI/CD
-
Have built and operated production infrastructure at a startup or larger company
-
Learn new solutions and technologies quickly
-
Improve and research infrastructure solutions continuously
-
Based in or willing to relocate to San Francisco to work in person at the hacker house
-
Willingness to work startup hours in a 996-style environment (9am–9pm, six days a week)
-
Quick learner who grasps products and systems fast
-
Experience building for performance and reliability at scale
-
Research and product focus mindset
-
High ownership mentality
-
Startup-minded operator who prioritizes learning and growth over work-life balance
-
GPU infrastructure experience in production
-
First infra hire at a startup
-
Background at an infrastructure company
-
Infra scope limited to model training pipelines only
-
20+ years of experience with a slow-moving, process-heavy background
-
Prioritizes work-life balance as a primary requirement
-
No production infra ownership
-
Sole infra owner with full-stack ownership from day one, directly in the critical path of live customer traffic.
-
Well-funded seed-stage company with strong early traction and an experienced backer base.
-
Significant equity, housing and food provided at the SF hacker house, visa sponsorship, laundry and cleaning, company off-sites, infinite DoorDash, and health & dental.
-
Location: San Francisco, CA
-
Work policy: In person (SF hacker house); 996 pace — 9am–9pm, six days a week
-
Compensation: $150,000–$250,000 + up to 1% equity
-
Visa sponsorship: Available (H-1B, O-1, OPT)
-
Employment type: Full-time