Staff Machine Learning Engineer, Platform (MLOPS)
Quick Summary
training or inference pipelines, model serving infrastructure, model registry and versioning, or production monitoring and alerting. Demonstrated owner
Credit Acceptance is proud to be an award-winning company recognized both locally and nationally across multiple workplace categories. Our world-class culture is shaped by dedicated team members who are driven to succeed as professionals individually and together as a team. Backed by a strong product, exceptional people, and a stable financial foundation, we’ve grown into a leading provider of used and new car financing across the country.
Our Engineering and Analytics Team Members utilize the latest technology to develop, monitor, and maintain complex practices that help optimize our success. Our Team Members value being challenged, are encouraged to express their ideas, and have the flexibility to enjoy work life balance. We build intrinsic value by partnering with all functions of our business to support their success and make strategic business decisions. We focus on professional development and continuous improvement while enjoying a casual work environment and Great Place to Work culture!
This is an operations and infrastructure role, not a modeling role. Success is measured by what stays up, what deploys safely, what is measurable in production, and what the platform costs to run.
- Customer Empathy: Customer Empathy is the ability to understand the perspectives, pain points, and experiences of customers. It involves actively putting oneself in the customer's shoes, comprehending their needs and challenges, and using that understanding to provide a better, more customer-centric experience.
- Engineering Excellence: Engineering Excellence is about bringing great craftsmanship and thought leadership to deliver an outstanding product that delights customers and solves for the business. This involves the pursuit and achievement of high standards, best practices, innovation, and superior solutions.
- One Team: A One Team mindset refers to a collaborative approach across the organization, where individuals work together seamlessly, without boundaries, as a single, cohesive team. Shared goals, open communication and mutual support create a sense of collective purpose. This enables teams to navigate challenges and pursue shared objectives more effectively.
- Owner's Mindset: Owner's Mindset involves adopting a set of behaviors that reflect a sense of responsibility, accountability, strategic thinking, and a proactive approach to managing your domain. As an owner, you understand the business and your domain(s) deeply and solve for the right outcome for the domain(s) and the business.
Requirements
~2 min read- Bachelor's degree in Computer Science, Engineering, Statistics or a relevant technical field with at least 7 years of relevant experience, or a Master's degree in one of those fields with at least 5 years of relevant experience.
- 5+ years building and operating production ML or AI systems, with direct ownership of at least two of: training or inference pipelines, model serving infrastructure, model registry and versioning, or production monitoring and alerting.
- Demonstrated ownership of a production ML or AI service through its full operating life: deployed it, monitored it, was paged for it, diagnosed a failure, and shipped the fix.
- Strong Python and SQL, with production-quality engineering practice: version control, testing, code review, and CI/CD applied to ML workloads rather than only to application code.
- Hands-on experience with a cloud ML platform in production. AWS and Databricks strongly preferred, including model serving, job orchestration, and a model registry or experiment tracking system such as MLflow.
- Working knowledge of how LLM and GenAI workloads differ operationally from traditional ML: token cost and latency behavior, caching and batching, non-deterministic output.
- Experience running LLM or agent applications behind a gateway or proxy layer, and able to build best practices for model routing and fallback, credential and key management, rate limiting and budget enforcement.
- Working knowledge of tool-calling architecture for agents, including the Model Context Protocol: what an MCP server is, how tools are scoped and authorized, and why tool access is brokered through a gateway rather than granted directly.
- Experience with containerization and infrastructure as code.
- Ability to communicate technical and non-technical trade-offs clearly in writing to an audience that includes both engineers and non-engineers.
Nice to Have
~1 min read- Production experience with GPU-backed model serving, including autoscaling behavior under concurrent load, cold-start management and cost control.
- Experience with OpenTelemetry and an enterprise observability platform such as Dynatrace, including instrumenting AI workloads rather than only conventional services.
- Experience with model serving efficiency techniques: quantization, parameter-efficient fine-tuning, distillation or inference optimization.
- Experience with data governance and access control on a lakehouse platform, such as Databricks Unity Catalog.
- Experience operating AI systems in a regulated industry, particularly financial services, including auditability, retention and access-control requirements.
- Experience with agentic or multi-step AI systems in production, including tracing and debugging multi-turn behavior.
- Hands-on experience with a managed agent or tool gateway, such as AWS Bedrock AgentCore Gateway, and with running MCP servers in a governed environment.
- Familiarity with emerging agent interoperability and agent identity standards, including Agent2Agent (A2A) style agent-to-agent delegation, agent discovery and capability advertisement, and workload identity for non-human actors.
- Ability to communicate complex technical information, both verbal and written, to all levels, including senior leadership.
- Ability to solve problems at the source by offering simple, working solutions.
- Responds promptly and effectively to resolve incidents, tasks, and projects.
- Demonstrated ability and motivation to teach others.
- Ability to gain the trust of others and build solid relationships across and vertically throughout the organization.
- Effectively prioritize and execute tasks in a high-pressure environment.
What We Offer
~1 min readTo be successful in this role, Team Members need to be:
- Positive by maintaining resiliency and focusing on solutions
- Respectful by collaborating and actively listening
- Insightful by cultivating innovation, accumulating business and role specific knowledge, demonstrating self-awareness and making quality decisions
- Direct by effectively communicating and conveying courage
- Earnest by taking accountability, applying feedback and effectively planning and priority setting
- Remain compliant with our policies processes and legal guidelines
- All other duties as assigned
- Attendance as required by department
Location & Eligibility
Listing Details
- Posted
- October 1, 2026
- First seen
- October 1, 2026
- Last seen
- October 4, 2026
Posting Health
- Days active
- 1
- Repost count
- 0
- Trust Level
- 65%
- Scored at
- October 3, 2026
Signal breakdown
Similar Staff Machine Learning Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.