Senior Applied ML Researcher
5 days ago
Cupertino, California, United States
Apple
Full-time
Free with email or Google
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
Free with email or Google
By continuing, you agree to our Terms & Privacy Policy.
We are seeking a Senior Applied ML Researcher to design, train, and deploy state-of-the-art models for visual and audio understanding. You will work on challenging problems at the intersection of computer vision, audio signal processing, and multimodal learning, enabling intelligent systems that can see, hear, and reason about the world. You will collaborate closely with research scientists, engineers, and product teams to find novel applications of Deep Machine Learning capabilities to assist our creative user base. Your mission is to elevate the workflows of millions of creators by combining generative AI with Apple’s human-centered design principles.
Design and train deep neural networks for video, image, audio, and audio-visual tasks. Build models for audio-visual representation learning, cross-modal alignment, and fusion. Develop solutions for tasks such as: Video understanding and temporal modeling. Audio-visual event detection. Speech, sound, and scene understanding. Multimodal classification, detection, and localization.
4+ years of experience in deep learning or machine learning engineering\\nstrong expertise in deep neural networks and modern training workflows\\n8 years + Hands-on experience with computer vision and/or audio modeling\\nProficiency in Python and deep learning frameworks (PyTorch preferred)\\nSolid understanding of linear algebra, probability, and optimization\\nAbility to build intuition from problem statement and translate to dataset requirement, neural network design and loss functions
PhD in computer science, machine learning, or a related field, or equivalent practical experience.\\nPublications in top-tier ML conferences (e.g., NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, etc.)\\nExperience with self-supervised or foundation model pre-training\\nOpen-source contributions in vision, audio, or multimodal AI\\nBonus: Experience with Objective-C and/or Swift for on-device deployment
Design and train deep neural networks for video, image, audio, and audio-visual tasks. Build models for audio-visual representation learning, cross-modal alignment, and fusion. Develop solutions for tasks such as: Video understanding and temporal modeling. Audio-visual event detection. Speech, sound, and scene understanding. Multimodal classification, detection, and localization.
4+ years of experience in deep learning or machine learning engineering\\nstrong expertise in deep neural networks and modern training workflows\\n8 years + Hands-on experience with computer vision and/or audio modeling\\nProficiency in Python and deep learning frameworks (PyTorch preferred)\\nSolid understanding of linear algebra, probability, and optimization\\nAbility to build intuition from problem statement and translate to dataset requirement, neural network design and loss functions
PhD in computer science, machine learning, or a related field, or equivalent practical experience.\\nPublications in top-tier ML conferences (e.g., NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, etc.)\\nExperience with self-supervised or foundation model pre-training\\nOpen-source contributions in vision, audio, or multimodal AI\\nBonus: Experience with Objective-C and/or Swift for on-device deployment