Software Engineer, Triage Services

23 hours ago

Cupertino, California, United States Apple Full-time
The System Triage Services team builds mission-critical applications that help detect, analyze, and classify low-level software crashes across all Apple platforms. You will design and build the next generation of AI-powered triage solutions, which serves as the first line of defense for all technologies developed by Core OS.\\n\\nThis role blends distributed systems engineering with hands-on agent development, tackling problems in reliability, scale, and system design alongside applied AI. If you want to understand operating systems at a deep level and are motivated by shipping work with measurable, visible impact, this role is for you\\n

You will design and build agentic systems that transform how Core OS triages software bugs and features at scale — building production-grade infrastructure that runs AI/ML workloads reliably and improves over time through rigorous evaluation and feedback loops. The goal is to make autonomous triage a dependable, first-class part of how Core OS ships and debugs software. You"ll work across the build and integration lifecycle, agent orchestration, and the infrastructure that connects them, partnering with project managers and engineering teams across Software, Hardware, and Silicon groups to turn complex triage requirements into resilient systems.\\n

Design, build, and maintain AI agent workflows that automate low-level software triage and root-cause analysis routing at scale.\\nEvaluate and improve agent reliability — including prompt design, tool-calling accuracy, guardrails, and regression testing.\\nInstrument agent pipelines end-to-end so failures, drift, and low-confidence routing decisions surface before reaching engineers.\\nCollaborate with Core OS software teams to identify automation opportunities and refine agent behavior through feedback and evaluation.\\nOwn the full lifecycle of production agents — from prototype to hardened service — including on-call ownership of triage-critical infrastructure, and mentor other engineers on agent design and evaluation.

Programming experience in Python\\nExperience building scalable data platforms and distributed systems.\\nExperience with LLM and agent development.\\nKnowledge of RAG architectures, embedding generation, vector databases, and AI data preparation for agentic workflows.\\nFamiliarity with eval frameworks for agent quality (offline benchmarks, drift detection, confidence scoring).\\nBachelor"s degree in Computer Science, Data Engineering, or a related field.\\n

Master"s degree in Computer Science, Data Engineering, or a related field.\\n1+ year developing production-grade agentic systems.\\nStrong collaboration and communication skills working across multi-domain software engineering teams.\\nExperience with large-scale telemetry or observability systems feeding automated decision-making.\\nExperience shipping tools adopted by engineering teams beyond your own.\\nBackground in kernel or OS-level debugging, crash analysis, or root-cause investigation.