Identify, align stakeholders on, and implement improvements to our processes, codebases, and architecture.
You are an exceptional SRE / DevOps engineer who is passionate about reliability, observability, clean automation, robust infrastructure architecture, and production stability. You view generative AI as a pragmatic utility to accelerate engineering and operational velocity.
You understand that technology best practices and patterns are always shifting, and you love evaluating whether emerging frameworks, tools, or practices should evolve our roadmap.
You have experience delivering flexible, highly available platforms that provide meaningful reliability and operational insights.
You possess a security-oriented mindset and deeply understand the security boundaries required when exposing infrastructure and automation tooling.
You treat everyone with empathy and respect.
You are a strong communicator, effectively clarifying technical needs, reliability trade-offs, and operational impacts for cross-functional and non-technical partners.
You excel at balancing rapid responses to real business and operational needs with foundational, long-term architectural and reliability stability.
14+ years of experience building, operating, and optimizing large-scale distributed systems, cloud infrastructure, and production platforms.
Deep hands-on expertise with Kubernetes (cluster architecture, networking, scaling, security, and day-2 operations).
Strong experience with Ansible (or equivalent configuration management) and Terraform (or equivalent Infrastructure as Code tools) for infrastructure automation, consistency, provisioning, and managing cloud and infrastructure resources.
Proven ownership of modern CI/CD pipelines, deployment strategies, and release engineering practices.
Extensive experience designing and operating observability platforms, with strong proficiency in Grafana (dashboards, alerting, and integration with metrics/logging/tracing systems).
Solid scripting and automation skills (Python, Bash, or equivalent) plus Infrastructure as Code practices.
Demonstrated ability to establish and drive SRE practices including SLIs/SLOs, error budgets, incident management, and toil reduction.
Experience using advanced coding/ops assistants extensively for day-to-day automation, infrastructure code, testing, and the ability to establish team standards around their safe and effective use.
Experience designing systems with appropriate human oversight for automated or agentic operational workflows.
2+ years of hands-on experience integrating generative AI/LLM components or advanced automation frameworks into production infrastructure or operational systems.
Familiarity with advanced Kubernetes operators, service meshes, or multi-cluster management.
Experience implementing tracing frameworks and advanced observability for complex distributed systems.
Knowledge of event-driven architectures, time-series data, or high-cardinality metrics environments.
Exposure to IoT telemetry or similar high-volume data streams.
Solar industry experience
We offer a competitive total compensation package that includes monthly health insurance premiums, bonuses and long-term stock options for every employee
We love to lift each other up through company-wide slack channels such as #puppiesandpets, #omnidian-wellness, #praiseandbooms and #sustainablefuture
We are a passionate, mission driven team that believes in collaboration, mutual respect and trust. For examples, come Discover our Story!
We mentor and invest in our employees and prioritize them for future opportunities. Check out our Instagram reels to see a few career journey examples
Internal candidates: Check out our advice on Internal Transfer: Job Application Process
We’re a fast-growing startup, which means we’re constantly reinventing processes, adding new products, and asking people to use all of their skills and talents. That means there’s gonna be a lot of opportunities for you to grow, which also means you will likely be stretched in ways you’ve never experienced in a job before. If you are resilient, determined, and not afraid of a big challenge, come apply.