Blue
Blue1mo ago
New

#700 - Incident Operations & Communications Engineer

Latam 2Remotemid
OtherCommunications Engineer
0 views0 saves0 applied

Quick Summary

Key Responsibilities

● Incident Response & Triage○ Respond to real-time alerts and determine whether they represent valid incidents.○ Initiate and participate in incident response workflows.

Technical Tools
OtherCommunications Engineer

BlueCloud is a Snowflake Elite Partner and the 2026 CoCo Catalyst Snowflake Partner of the Year. We help enterprise organizations move from fragmented legacy systems to unified, AI-ready Snowflake platforms — delivering data migration, engineering, governance, BI & analytics, and AI/ML solutions 40–50% faster than traditional approaches.

With 450+ Snowflake consultants, 200+ enterprise transformations under our belt, and a 100% Snowflake focus, we combine advisory-led thinking with AI-powered accelerators to turn months of work into weeks of results. Our clients span Financial Services, Healthcare & Life Sciences, Retail, Manufacturing, Energy, and more — and the outcomes speak for themselves: 97% faster reports, 40% fraud reduction, $1.5M in client savings, and 10× client growth.

We don't just strategize — we execute.


Position Summary:
We're seeking a skilled and detail-oriented Incident Operations & Communications Engineer to manage real-time incident response, external communications, and coordination across our production systems. This role sits at the intersection of Software Engineering, Site Reliability Engineering (SRE), Incident Response, and Partner Operations, and is ideal for someone who thrives in a fast-paced, high-stakes environment requiring rapid decision-making under ambiguity. You will play a critical role in the incident lifecycle, including impact assessment, SLA-driven communications, and cross functional coordination with Engineering and Technical Account Managers (TAMs), ensuring timely, accurate, and consistent communication with merchants and partners. This position is a key component to support global operations and enable scalable External Incident Comms as we expand internationally.
 

Key Responsibilities:
Incident Response & Triage
○ Respond to real-time alerts and determine whether they represent valid incidents.
○ Initiate and participate in incident response workflows.
○ Collaborate with Incident Commanders and different Stakeholders.
○ Manage multiple concurrent incidents while maintaining clarity and prioritization.

Incident Lifecycle Coordination
○ Act as a central coordination point between:
■ Engineering teams
■ Incident Commanders
■ TAMs and partner-facing teams
○ Support handoffs across time zones as part of the FTS model.
○ Contribute to incident closure, RCA inputs, and reporting workflows, including
ownership of SLA merchants and distribution of external SLA Reports.

Impact Assessment
○ Identify affected merchants and partners using dashboards, alerts, and system
signals.
○ Make rapid decisions with incomplete or evolving data.
○ Evaluate incident severity and determine communication requirements based on
SLA commitments.
○ Continuously update impact scope as incidents evolve.

External Communications
○ Own end-to-end communication lifecycle with merchants and partners:
■ Initial notifications within strict SLA windows
■ Ongoing updates aligned with severity-based cadence
■ Final resolution communications
○ Tailor messaging based on merchant-specific requirements and communication
rules.
○ Coordinate approvals with stakeholders (TAMs, Engineering, Leadership) when
required.
○ Ensure communication quality, clarity, and consistency across all outputs.

Status Page Management
○ Create and maintain incident entries on merchant-facing status pages.
○ Align internal incident state with external communication.
○ Maintain update cadence based on incident severity (e.g., SEV0, SEV1).
○ Ensure compliance with publishing guidelines and approval processes.

● Operational Scaling & Process Improvements
○ Identify opportunities to reduce manual toil across the incident lifecycle.
○ Contribute to development of:
■ Standardized communication workflows
■ Automation tools and internal systems
■ Improved impact assessment methodologies
○ Collaborate with Engineering and TAMs (NA & EU) to improve:
■ Observability and monitoring
■ Communication decision frameworks
○ Support rollout of new roles and capabilities (e.g., Scribe) within the incident
lifecycle.
○ When not on active rotation, contribute directly to core reliability-loop tooling — Scribe, Project Nemo (Deepdive), and Project Elcano — (Python/Kotlin), with hands-on feature ownership, not just identifying improvement opportunities.

Qualifications:
Experience:
○ 4/5+ years of experience in Software Engineering, Incident Operations, DevOps/SRE, or similar roles.
○ Hands-on experience with Python (Kotlin and frontend exposure as a nice to have)
○ Experience working in on-call environments with SLA-driven responsibilities.
○ Familiarity with observability tooling and dashboards and DevOps/SRE concepts
○ Proven ability to operate in high-pressure, real-time incident scenarios

Technical Background:
○ Strong understanding of distributed systems and production environments.
○ Experience with:
■ Monitoring and alerting systems (e.g., Datadog, Chronosphere, etc.)
■ Incident management tools (e.g., PagerDuty, Rootly, Slack workflows)

Familiarity with:
■ APIs and system integrations
■ Understanding of software development lifecycle (SDLC) and production reliability.

● Soft Skills:
○ Strong operational judgment and ability to make decisions under uncertainty.
○ Excellent written and verbal communication skills, especially in external-facing contexts.
○ Ability to manage multiple priorities simultaneously in a high-pressure environment.
○ Highly organized and detail-oriented, with a strong ownership mindset.
○ Strong cross-functional collaboration skills, especially with Engineering and partner teams.
○ Proactive mindset with focus on continuous improvement and scalability
We are looking for someone who is experienced in building lasting relationships and is passionate about making meaningful contributions to our team. Don't be mistaken, this is a challenging career path but also highly rewarding. Are you up for the challenge? If so, stop reading and start applying.

Location & Eligibility

Where is the job
Worldwide
Fully remote, anywhere in the world
Who can apply
Same as job location

Listing Details

Posted
August 12, 2026
First seen
September 16, 2026
Last seen
September 16, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
30%
Scored at
September 16, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Blue
Blue
lever

Gopher

Employees
350
Founded
2018
View company profile
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Blue#700 - Incident Operations & Communications Engineer