Staff Machine Learning Engineer
TLDR
Lead production-grade ML services across cloud and edge, bridging research and software engineering.
-
Lead the Machine Learning Engineering team in designing, building, deploying, and operating production-grade machine learning services.
-
Design scalable backend architectures for serving machine learning models through secure and maintainable APIs.
-
Own the production lifecycle of machine learning models, including deployment, monitoring, versioning, validation, rollback, and retirement.
-
Build and maintain automated ML pipelines covering model packaging, testing, deployment, retraining, and release.
-
Establish best practices for model governance, including version management, experiment tracking, dataset versioning, lineage, and auditability.
-
Work closely with the AI Research team to productionize validated research outcomes while providing engineering feedback during research.
-
Collaborate with Full Stack Engineering teams to deliver well-documented, reliable APIs that enable seamless product integration.
-
Ensure deployed models meet requirements for scalability, reliability, latency, observability, and security.
-
Drive engineering standards, technical direction, and continuous improvement across the ML Engineering team.
-
Mentor engineers, conduct technical reviews, and foster a culture of engineering excellence.
-
Evaluate emerging technologies and tooling to continuously improve the ML engineering platform.
-
Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, or a related discipline, or equivalent practical experience.
-
Typically 8+ years of software engineering experience, including significant experience building production machine learning systems.
-
Proven experience leading engineering teams or technical initiatives involving machine learning platforms.
-
Strong experience deploying and operating machine learning models in cloud-based production environments.
-
Experience deploying machine learning solutions to edge devices is highly desirable.
-
Strong understanding of the complete machine learning lifecycle, including:
-
Model packaging and deployment
-
Continuous retraining strategies
-
Model validation and qualification
-
Model versioning
-
Dataset versioning and lineage
-
Experiment tracking
-
Performance monitoring
-
Drift detection
-
Model rollback strategies
-
Compliance and reproducibility requirements
-
Strong backend software engineering background with experience designing scalable RESTful APIs and service-oriented architectures.
-
Experience designing scalable distributed systems and microservices.
-
Experience implementing CI/CD pipelines for machine learning systems (MLOps).
-
Strong understanding of cloud platforms such as AWS, Azure, or Google Cloud.
-
Excellent problem-solving skills with the ability to simplify complex technical challenges.
-
Excellent communication skills with the ability to collaborate across research, engineering, product, and leadership teams.
-
Experience with Amazon SageMaker or equivalent enterprise ML platforms.
-
Experience with MLFlow, Kubeflow, Vertex AI, Azure ML, or similar MLOps platforms.
-
Experience with feature stores and model registries.
-
Knowledge of GPU-based inference and optimization techniques.
-
Experience deploying large language models (LLMs), computer vision systems, or multimodal AI applications.
-
Experience with edge AI optimization frameworks such as TensorRT, ONNX Runtime, OpenVINO, or similar technologies.
-
Hands-on experience with containerization technologies such as Docker and Kubernetes.
-
Experience working in product-focused or fast-growing technology companies.
-
Leadership Expectations
-
Provide technical leadership and direction for the Machine Learning Engineering function.
-
Establish engineering standards, operational processes, and best practices for production AI systems.
-
Influence architectural decisions across engineering teams.
-
Develop team capability through mentoring, coaching, and technical leadership.
-
Build strong partnerships with AI Research, Full Stack Engineering, DevOps, Product Management, and Technical Project Management teams.
-
Drive execution while balancing engineering quality, operational reliability, and business priorities.
Robotic Assistance Devices builds advanced robotic solutions designed to improve operational efficiency across multiple industries. Our technology caters to businesses seeking to streamline processes and enhance productivity through automation.