Senior Data Analyst

3 days ago

El Segundo CA, Los Angeles County, CA; California, United States Rightsline Inc Full-time

Reporting to the SVP Engineering, the Senior Data Analyst – LLMOps will work with the Application Development team to own the observability, evaluation, and quality measurement of Rightsline’s LLM-powered features. This is a senior individual-contributor role in which you will instrument the LLM stack end to end — traces, prompts, tool calls, latency, token spend, and output quality — build the evaluation suites that tell us whether a change made the product better or worse, and turn that telemetry into decisions the engineering and product teams act on. You will work across a SaaS application, a wide range of third-party tools, an AWS-hosted infrastructure, and an increasingly agentic development workflow. The position is based in North America, has no direct reports, and partners closely with the wider engineering, product, and data teams.

Mission & Impact

The Senior Data Analyst – LLMOps makes the behavior of Rightsline’s AI features measurable. Quality questions about prompts, retrieval, and agent workflows are too often answered with intuition and spot checks; you will replace that with tracing, offline and online evaluations, regression suites, and dashboards that show quality, cost, and latency trends over time. You will be working in a collaborative environment with a team who shares your passion for building something impactful, and success in the role looks like engineers and product managers reaching for your evaluation results and dashboards before they ship — and trusting what they see.

What You Will Own

This role’s work spans five areas, each described below.

Observability and Telemetry

  • Instrument Rightsline’s LLM and agent workflows end to end with tracing and structured logging, using tools such as LangSmith, Langfuse, or equivalent

  • Define and maintain the metrics that describe AI system health — latency, token and cost per request, error and fallback rates, tool-call success, and retrieval hit rates

  • Build and maintain dashboards and alerting so quality and cost regressions surface in hours rather than in customer tickets

Evaluation and Quality

  • Design and maintain offline evaluation suites — golden datasets, scoring rubrics, and LLM-as-judge grading — for prompts, retrieval pipelines, and agent flows

  • Run online evaluations and A/B experiments on production traffic, and report results with enough rigor to support a ship / no-ship decision

  • Track quality across model, prompt, and retrieval changes, and maintain the benchmark history that shows whether the product is actually improving

  • Partner with engineers on human review and annotation workflows, including labeling guidelines and inter-rater agreement

Data Analysis and Reporting

  • Build and maintain the pipelines that move trace, evaluation, and usage telemetry into the analytics stack using SQL and Python

  • Analyze how customers actually use AI features — adoption, drop-off, and failure patterns — and translate findings into prioritized recommendations

  • Produce recurring reporting on AI quality, cost, and usage for engineering, product, and leadership audiences

Cost, Reliability, and Troubleshooting

  • Monitor token spend and model usage, and identify where prompt, model, or caching changes cut cost without hurting quality

  • Investigate quality incidents — bad or unsafe outputs, hallucinations, tool failures, latency spikes — from trace to root cause

  • Follow and help extend Rightsline’s development best practices and standards in the pipelines, queries, and evaluation code you ship

Team Contribution

  • Attend and contribute to team meetings, and present telemetry and evaluation findings in a way non-specialists can act on

  • Raise the team’s LLMOps literacy — help engineers instrument their own features, write their own evals, and read the dashboards

What You Will Bring

  • 5+ years in data analytics, analytics engineering, or data science, including hands-on work with LLM-based systems in production

  • Strong SQL and Python (including pandas and notebook-base