Bear Robotics
Quick Summary
Platform Architecture & Design Design and build the end-to-end platform for robot data, model training, evaluation, versioning, release and deployment. Establish model and dataset registries,
Strong experience building production ML platforms, data-intensive systems or distributed training infrastructure used by research or engineering teams.
Responsibilities
~1 min read- →None.
- Design and build the end-to-end platform for robot data, model training, evaluation, versioning, release and deployment.
- Establish model and dataset registries, artefact versioning, configuration management and traceability from an on-robot result back to code, data and parameters.
- Create reliable ingestion and processing pipelines for multimodal robot data, including video, depth, proprioception, actions, language, events and operational metadata.
- Develop tools for dataset discovery, curation, labelling, lineage, quality control, balancing, replay and governance at increasing scale.
- Build efficient distributed training and batch-evaluation systems that make effective use of GPU clusters and support reproducible experimentation.
- Create automated evaluation, regression and release workflows with clear promotion criteria and safe rollback paths.
- Develop deployment tooling for edge and robot compute, including packaging, optimisation, staged rollout, monitoring and fleet-level model management.
- Improve platform reliability, observability, security, cost efficiency and developer experience; define service-level expectations for critical workflows.
- Partner closely with research, robot learning, perception, simulation, infrastructure and field teams to convert recurring friction into reusable platform capability.
- Lead platform architecture decisions, write clear technical proposals and mentor engineers in scalable ML-systems practices.
- Performs other duties or takes on specialized responsibilities as assigned.
Requirements
~2 min read- Strong experience building production ML platforms, data-intensive systems or distributed training infrastructure used by research or engineering teams.
- Excellent Python and software-engineering skills, with experience designing maintainable services, APIs, libraries and workflow tooling.
- Practical experience with cloud infrastructure, containers and orchestration, Docker and infrastructure-as-code.
- Experience with modern data and workflow technologies, distributed storage and compute, and the operational realities of large multimodal datasets.
- Knowledge of GPU workloads, PyTorch-based training, profiling, scheduling and performance or cost optimisation.
- Strong understanding of ML lifecycle concerns: reproducibility, lineage, model registries, evaluation, deployment, monitoring and rollback.
- Ability to translate diverse researcher needs into coherent abstractions without over-engineering early solutions.
- Track record of owning technically ambiguous platform work and operating production systems reliably.
- Robotics or autonomous-systems data, especially high-bandwidth time-series and synchronised multimodal logs.
- Distributed training at scale, cluster scheduling, checkpointing, fault tolerance and experiment orchestration.
- Data engines for imitation learning, reinforcement learning, active learning or continual learning.
- Edge inference, model compilation and optimisation using ONNX, TensorRT, CUDA or related technologies.
- Hybrid cloud and on-premise compute, fleet management, intermittent connectivity and remote deployment.
- Security, privacy and access controls for sensitive operational data.
- Developer platforms, internal tooling and measurable improvements to researcher productivity.
The physical demands described here are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.
- Prolonged periods of sitting/standing at a desk and working on a computer. The employee routinely is required to sit; stand, walk; talk and hear; use hands to keyboard.
- Specific vision abilities required by this job include close vision, color vision, peripheral vision, depth perception, and ability to adjust focus.
- Ability to lift 30 lbs.
- Occasional travel to partner or pilot sites may be required.
- A degree in computer science, engineering or a related field, or equivalent practical experience.
- At least three years' relevant work experience.
Location & Eligibility
Listing Details
- First seen
- September 9, 2026
- Last seen
- September 9, 2026
Posting Health
- Days active
- 0
- Repost count
- 1
- Trust Level
- 52%
- Scored at
- September 9, 2026
Signal breakdown
Please let Bear Robotics know you found this job on Jobera.
3 other jobs at Bear Robotics
View all →Explore open roles at Bear Robotics.
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.
