Sr Forward Deployment Engineer
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
By continuing, you agree to our Terms & Privacy Policy.
Senior Forward Deployed GenAI Engineer
Hybrid Opportunity- Chicago (3 days a week at clients office)
Re-location provided.
About Avanta
Avanta helps ambitious small and mid-market companies build AI intelligence that changes how work gets done—in weeks, not quarters. We design and deliver practical AI systems from model selection through production deployment, and we build them so the customer can understand, operate, and own the outcome.
Our work spans agentic solutions, AI enablement, AI-native software delivery, business transformation, small language models, strategy, governance, adoption, measurement, and executive advisory.
Role Summary
You will independently design and deliver production generative AI applications while balancing customer value, technical quality, security, cost, adoption, and operational ownership. You will mentor associate engineers and serve as a trusted technical partner to customer teams.
What You Will Do
- Lead the design and implementation of production-grade generative AI applications and customer workstreams.
- Build RAG systems across structured, semi-structured, and unstructured data, including authorization-aware retrieval and citations.
- Design agentic workflows coordinating multiple tools, services, business rules, and human approval boundaries.
- Evaluate proprietary, open-weight, multimodal, embedding, reranking, and specialized models for specific business needs.
- Benchmark quality, accuracy, latency, throughput, cost, safety, and operational complexity before recommending a model or architecture.
- Create golden datasets, scoring rubrics, regression tests, and repeatable evaluation pipelines.
- Diagnose failures across ingestion, parsing, chunking, embeddings, retrieval, reranking, prompts, tool execution, orchestration, and output generation.
- Implement human review, approval, escalation, and rollback mechanisms for higher-impact actions.
- Design for multitenancy, resiliency, availability, graceful degradation, and defined service-level objectives.
- Integrate AI applications with enterprise systems, data platforms, workflow tools, identity services, and security controls.
- Build production monitoring for model quality, retrieval effectiveness, tool execution, latency, token usage, cost, errors, and user behavior.
- Improve systems through caching, batching, context reduction, prompt optimization, model routing, provider abstraction, and architectural changes.
- Build deployment pipelines, infrastructure automation, automated testing, and release-management practices.
- Support production incidents, lead root-cause analysis for owned workstreams, and implement durable corrective actions.
- Convert customer implementations into templates, connectors, evaluation packs, runbooks, accelerators, and reference architectures.
- Mentor associate engineers and provide code, design, and customer-communication feedback.
Customer-Facing Responsibilities
- You will lead technical discovery with customer engineers, business owners, security teams, operators, and data stakeholders.
- You'll translate business problems into measurable AI use cases and technical requirements, facilitate design workshops, create implementation plans, coordinate across customer teams, communicate risks, and help customer engineers understand and operate the delivered solution.
- You may also support pre-sales technical discovery when needed, while remaining accountable for hands-on delivery .
Technical Responsibilities
- Design cloud-native services using API, distributed-system, asynchronous, and event-driven patterns.
- Apply hybrid retrieval, metadata filtering, query rewriting, reranking, retrieval evaluation, and graph-based retrieval where appropriate.
- Design agent state, memory, tool contracts, authorization, failure recovery, and deterministic controls.
- Establish prompt/configuration versioning, automated regression testing, experiment tracking, and release gates.
- Implement identity-aware retrieval and least-privilege access.
- Protect systems from prompt injection, data leakage, unsafe tool use, malicious input, and invalid structured output.
- Apply containers, Kubernetes or serverless patterns, infrastructure as code, and automated CI/CD.
- Trace and observe the complete application path across models, retrieval, agents, tools, data, and infrastructure.
- Optimize model and application cost, latency, throughput, and reliability.
- Define runbooks, incident procedures, support boundaries, SLOs, and ownership