Senior Software Engineer, ML/AI Platform
Attentive· United States·
About the Role
We’re looking for a self-motivated, highly driven Senior Software Engineer to join Attentive’s Machine Learning Platform team. As a hands-on individual contributor, you’ll build and operate the platform capabilities that enable ML engineers and data scientists to develop, train, evaluate, deploy, and serve models reliably at scale.
Your work will span the ML lifecycle—including data and feature access, training infrastructure, model lifecycle management, batch and real-time inference, and platform observability. You’ll own well-scoped projects from design and implementation through production operations, partnering with Staff engineers and platform users to translate broader architectural direction into reliable, self-service capabilities.
This is an opportunity to join a growing ML Platform team and directly improve how quickly and safely Attentive ships AI- and ML-powered products.
What You’ll Accomplish
- Build and operate production-grade services and workflows for training, evaluating, deploying, and serving ML models.
- Implement improvements across ML data and feature access, training infrastructure, model lifecycle tooling, and batch and online inference.
- Develop scalable APIs, abstractions, and self-service tools that make ML engineers and data scientists more productive.
- Improve the reliability, observability, performance, and cost efficiency of ML platform components.
- Own projects through implementation, testing, rollout, monitoring, and ongoing production support.
- Contribute to technical designs and help break larger platform initiatives into deliverable milestones.
- Partner with ML, Data Science, Product Engineering, and Infrastructure teams to deliver AI and ML initiatives.
- Raise engineering quality through code reviews, documentation, operational best practices, and mentorship.
Your Expertise
- 5+ years building and operating production software or distributed systems, with meaningful experience in ML platform, MLOps, or ML infrastructure.
- Strong software engineering skills in Python, Java, or a comparable language, including experience building production services or developer-facing platforms.
- Experience in one or more areas of the ML lifecycle: training infrastructure, orchestration, feature platforms, model deployment, model lifecycle management, or inference.
- Experience with distributed data or compute technologies such as Spark, Ray, Kafka, or similar systems. Expertise in every listed technology is not required.
- Understanding of the design and operational tradeoffs between batch, streaming, online, and offline ML workloads.
- Experience running reliable workloads in a cloud environment using technologies such as Kubernetes, AWS, and infrastructure-as-code.
- Ability to diagnose performance, scalability, and reliability problems across application, data, and infrastructure layers.
- A track record of independently delivering complex projects while collaborating effectively with technical and business partners.
What We Use
- AWS, EKS, Kubernetes, Terraform, Helm, Istio, and Datadog
- Metaflow, MLflow, Argo, PyTorch, TensorFlow, and Hugging Face
- Python and Java services
- DynamoDB, Postgres, Redis, Kinesis, and other production data systems
- Spark, Ray, Kafka, and distributed batch and streaming systems
You'll get competitive perks and benefits, from health & wellness to equity, to help you bring your best self to work.
For US based applicants:
- The US base salary range for this full-time position is $180,000 - $250,000 annually + equity + benefits
- Our salary ranges are determined by role, level and location
#LI-EZ1
By applying for this position, your data will be processed as per Attentive's Privacy Policy.