True Zero is seeking a Datadog Observability Engineer to lead the design, implementation, and operation of enterprise-grade observability solutions built on the Datadog platform across secure, mission-critical, and healthcare environments. This is a senior technical role for an engineer who pairs deep, hands-on expertise in infrastructure monitoring, APM, log management, and synthetic testing with the judgment to guide teams, engage clients, and deliver production-grade, highly reliable observability capabilities.
The ideal candidate brings a track record of architecting and operating Datadog environments across large, distributed systems — including experience monitoring healthcare applications and infrastructure subject to HIPAA and other regulatory requirements, where uptime, data integrity, and patient-safety-critical alerting are paramount. Experience integrating endpoint and asset visibility from Tanium into the overall observability and security architecture is a strong plus, giving teams a unified view across infrastructure health, application performance, and endpoint risk.
Observability Architecture: Design, implement, and maintain end-to-end observability solutions using Datadog — infrastructure monitoring, APM/distributed tracing, log management, RUM, and synthetic monitoring — across cloud and on-premises environments.
Dashboard & Alerting Development: Build and maintain dashboards, SLOs, and intelligent alerting policies that surface actionable signal and reduce noise for on-call teams.
Healthcare Systems Monitoring: Implement and tune observability for healthcare applications and infrastructure (EHR/EMR platforms, HL7/FHIR interfaces, clinical systems) to protect uptime, data integrity, and compliance with HIPAA and related regulatory requirements.
Log & Metrics Pipeline Engineering: Configure log pipelines, custom metrics, and telemetry ingestion (via Datadog Agent, OpenTelemetry, and APIs) to ensure high-fidelity, cost-efficient data collection at scale.
Incident Response & On-Call Support: Partner with SRE and engineering teams to detect, triage, and resolve incidents; drive root-cause analysis and post-incident reviews using Datadog’s tooling.
Systems & Security Integration: Integrate the Datadog platform with existing ITSM, CI/CD, and security tooling; collaborate with security teams to incorporate endpoint and asset data from Tanium into the overall observability and security architecture for unified risk and health visibility.
Infrastructure as Code: Manage Datadog configuration (monitors, dashboards, SLOs) as code using Terraform or the Datadog API/CLI to ensure consistency, version control, and repeatable deployments.
Cost & Performance Optimization: Monitor and optimize Datadog usage, tagging strategy, and data retention to control cost while maintaining coverage and performance.
Governance & Compliance: Ensure observability practices align with regulatory and security requirements, including HIPAA, NIST SP 800-53, and Zero Trust principles.
Reporting: Deliver technical reports on system health, performance trends, and reliability posture with clear, actionable recommendations.
Client Collaboration: Engage with clients and stakeholders throughout the project lifecycle to align observability strategy with mission and business objectives.
Team Development: Coach and mentor junior engineers on observability best practices and Datadog platform capabilities.
Experience: Minimum of 5 years in observability, DevOps, SRE, or systems engineering, with substantial hands-on experience administering and architecting the Datadog platform.
Citizenship: U.S. citizenship required; must be willing to undergo a U.S. Government background investigation.
Datadog Expertise: Deep hands-on experience with Datadog APM, Infrastructure Monitoring, Log Management, Synthetic Monitoring, RUM, and Cloud Security Management.
Healthcare Experience: Experience monitoring and supporting healthcare applications and infrastructure (EHR/EMR, HL7/FHIR, clinical or patient-facing systems) in HIPAA-regulated environments is highly valued.
Endpoint & Security Integration: Familiarity with Tanium or similar endpoint management/visibility platforms, and experience incorporating endpoint telemetry into the overall observability and security architecture, is a strong plus.
Cloud & Infrastructure: Hands-on experience with major cloud platforms (AWS, Azure, GCP), containerized environments (Kubernetes, Docker), and hybrid/on-prem infrastructure.
Automation & IaC: Proficiency with Terraform and scripting/automation (Python, Bash) for configuration-as-code and pipeline automation.
Methodologies: Strong understanding of SRE principles, SLIs/SLOs/error budgets, incident management, and root-cause analysis.
Governance: Familiarity with NIST SP 800-53, HIPAA technical safeguards, and Zero Trust architecture principles.
Integration: Experience integrating observability tooling with ITSM platforms (ServiceNow, Jira), CI/CD pipelines, and SIEM/security tooling.
Communication: Excellent written and verbal skills to convey technical information to diverse audiences.
Education: Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field preferred; equivalent experience considered.