Software Engineer
3 days ago
ABOUT BASETEN
Baseten powers inference for the world's most dynamic AI companies, like OpenEvidence, Clay, Mirage, Gamma, Sourcegraph, Writer, Abridge, Bland, and Zed. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. With our recent $150M Series D funding, backed by investors including BOND, IVP, Spark Capital, Greylock, and Conviction, we're scaling our team to meet accelerating customer demand.
THE ROLE:
Baseten's Model Performance (MP) team is responsible for ensuring the models running on our platform are fast, reliable, and cost‑efficient. As part of this team, you'll focus on Model API's — the infrastructure powering our hosted API endpoints for the latest open‑source models. This work spans distributed systems, model serving, and developer experience. You'll join a small, high‑impact team operating at the intersection of product, model performance, and infra, helping to define how developers interact with AI models at scale.
RESPONSIBILITIES:
Design, build, and operate the Model APIs surface with focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling and multi-modal serving
Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, implement custom CUDA operators, tune memory allocation patterns for maximum throughput and optimize communication patterns across multi-GPU setups
Productionize performance improvements across runtimes with deep understanding of their internals: speculative decoding implementations, guided generation for structured outputs, custom scheduling and routing algorithms for high-performance serving
Build comprehensive benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware configurations
Productionize performance improvements across runtimes (e.g.TensorRT, TensorRT‑LLM): speculative decoding, quantization, batching, and KV‑cache reuse.
Instrument deep observability (metrics, traces, logs) and build repeatable benchmarks to measure speed, reliability, and quality.
Implement platform fundamentals: API versioning, validation, usage metering, quotas, and authentication.
Collaborate closely with other teams to deliver robust, developer‑friendly model serving experiences.
REQUIREMENTS:
3+ years experience building and operating distributed systems or large‑scale APIs.
Proven track record of owning low‑latency, reliable backend services (rate‑limiting, auth, quotas, metering, migrations).
Infra instincts with performance sensibilities: profiling, tracing, capacity planning, and SLO management.
Comfortable debugging complex systems, from runtime internals to GPU execution traces.
Strong written communication; able to produce clear design docs and collaborate across functions.
NICE TO HAVE:
Experience with LLM runtimes (vLLM, SGLang, TensorRT‑LLM) or contributions to open-source inference engines (vLLM, TensorRT-LLM, SGLang, TGI)
Knowledge of Kubernetes, service meshes, API gateways, or distributed scheduling.
Background in developer‑facing infrastructure or open‑source APIs.
We value infra‑leaning generalists who bring strong engineering fundamentals and curiosity. ML experience is a plus, but not required.
BENEFITS
Competitive compensation package.
This is a unique opportunity to be part of a rapidly growing startup in one of the most exciting engineering fields of our era.
An inclusive and supportive work culture that fosters learning and growth.
Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Apply now to embark on a rewarding journey in shaping the future of AI If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.
-
Software Engineer
2 days ago
San Francisco, California, United States Beacon Software Full time $120,000 - $250,000 per yearBeacon Software is a permanent capital holding company which acquires and grows essential businesses. We are a profitable series B+ firm that combines great technologists, operators and M&A professionals to accelerate the scale of the ambition of the dozens of businesses we own and operate. We are supported by capital from tier-1 venture capital, crossover,...
-
Software Engineer
4 days ago
San Francisco, California, United States Acquired Talent Ltd. Full time $200,000 - $225,000 per yearFounding Software Engineer / Ruby / Ruby on Rails / React / Postgres / AWS / Kubernetes / Software Engineer / APIsFounding Software EngineerSalary: $200,000 to 225,000Meaningful EquityLocation: San Francisco, 4 days a week in officeWe're on the lookout for a Key Founding Software Engineer (Ruby & React) for a revolutionary SaaS start-up that are helping SMBs...
-
Software Engineer III
2 days ago
San Mateo, California, United States Guidewire Software Full time $128,000 - $192,000SummaryGuidewire is widely recognized as the leading provider of software solutions for the property and casualty insurance industry. Our software empowers insurance companies to support their customers during their critical times – whether in response to natural disaster, an accident or emerging risks.At Guidewire, we offer a dynamic work...
-
Software Engineer III
3 days ago
San Mateo, California, United States Guidewire Software, Inc. Full time $128,000 - $192,000 per yearUnited States - San Mateo, CAProduct Strategy/Full time/HybridGuidewire is widely recognized as the leading provider of software solutions for the property and casualty insurance industry. Our software empowers insurance companies to support their customers during their critical times – whether in response to natural disaster, an accident or emerging...
-
Software Engineer
2 weeks ago
San Francisco, California, United States Crossing Hurdles Full timeCrossing Hurdlesis a global recruitment consultancy partnering with a fast-growing startup focused on empowering ESL learners. The startup has achieved a3.7x growth trajectoryand is targeting to exceed$3M+ in revenue within the next year.Role - Backend Software EngineerLocation -1 day per week in Palo Alto officeRequired Tech Stack -Python, FastAPI,...
-
Software Engineer
2 days ago
San Francisco, California, United States ExecutivePlacements Full time $150,000 - $250,000 per yearJoin to apply for theSoftware Engineer (Backend)role atColumnJoin to apply for theSoftware Engineer (Backend)role atColumn*About Column*For companies building financial technology and transforming the financial services space, the biggest bottleneck to their growth and innovation is often the underlying banks and infrastructure stack they rely on. We have...
-
Frontend Software Engineer
15 hours ago
San Francisco, California, United States PlayerZero, Inc. Full time $120,000 - $180,000 per yearSenior Software Engineer; Frontend EngineerAbout PlayerZeroPlayerZero is building a self-healing system for software that automates defect resolution and development. We are used by engineering and support teams to:autonomously debug problems in the software (technical support)fix issues directly in the codeprevent these problems from recurringPlayerZero is...
-
Robotics Software Engineer
4 days ago
San Francisco, California, United States Skild AI Full time $100,000 - $300,000 per year*We are hiring in San Francisco and Pittsburgh locationsCompany OverviewAt Skild AI, we are building the world's first general purpose robotic intelligence that is robust and adapts to unseen scenarios without failing. We believe massive scale through data-driven machine learning is the key to unlocking these capabilities for the widespread deployment of...
-
Software Engineer
10 hours ago
San Francisco, California, United States Attentive Full time $136,000 - $175,000 per yearAttentive is the AI-powered mobile marketing platform transforming the way brands personalize consumer engagement. Attentive enables marketers to craft tailored journeys for every subscriber, driving higher recurring revenue and maximizing campaign performance. Activating real-time data from multiple channels and advanced AI, the platform personalizes...
-
Software Engineer
2 days ago
San Francisco, California, United States Liberate Full time $60,000 - $72,000 per yearAbout Us:Liberate Innovations Inc. is a Series-B funded AI company focused on revolutionizing the insurance industry through cutting-edge technology solutions. We are looking for a skilled Senior Software Engineer to join our development team and help build scalable, reliable, and innovative AI-driven applications.You will join an amazing team of talented...