Even a simple configuration error can lead to significant unintended usage. At HighLevel, we aim to build robust systems that protect our customers and the platform from accidental surges and intentional abuse alike. This team creates the intelligent safeguards that ensure our automation remains safe and reliable for everyone.
We are looking for a Staff Engineer to architect the future of Trust, Controls & Safety. You will lead the technical direction of a real-time platform that secures critical workflows across HighLevel—including Signups, SaaS checkouts, User management, and global communication usage. This system processes 160M+ events and $4M+ in transactions daily, serving as the essential foundation for platform integrity.
The challenge is to build a highly available, low-latency control plane that balances strict enforcement with a seamless user experience. By integrating synchronous metering with asynchronous signal correlation, you will help us stay ahead of evolving fraud vectors while ensuring that legitimate business activities, from new user signups to complex SaaS billing cycles, proceed without friction.
As a Staff Engineer, you'll set the technical direction, own the hardest problems on the platform, mentor engineers across squads, and partner closely with EMs, PMs, Finance, Risk, Support and the channel teams. Getting this right means customers stop getting bills they didn't intend, and abuse stops being a cost the honest majority carries.
Database architecture: Redesign and optimize data models, queries, and indexing strategies across MongoDB, Firestore, and ElasticSearch
Search at scale: Own ElasticSearch reliability — ingestion, indexing, shard strategy, and query performance for billions of documents
Reliability & performance: Eliminate bottlenecks, improve replication health, and enforce predictable query and index latency
System design: Define data flow and integration boundaries between storage, cache, and APIs with clear contracts and fault isolation
Observability: Build visibility into every data path — metrics, traces, slow query logs, replication lag, and cluster health dashboards
Testing & quality: Make testing non-negotiable — enforce unit, integration, and load testing standards across all backend modules
Engineering culture: Drive RFCs, ADRs, and design reviews that push clarity and precision. Codify patterns that make good engineering repeatable
Hands-on leadership: Write code, design systems, and debug production issues. Lead by technical example, not by delegation
Mentorship & influence: Level up engineers around you — reviews that teach, feedback that sticks, and systems that outlive individuals
Heavyweight Engineering Pedigree: 8+ years of experience building high-scale, mission-critical backend systems where p99 regressions are unacceptable.
Distributed Systems Mastery: Deep expertise in concurrency, event-driven architectures, and consistency primitives (CAS, atomic dry-runs, idempotency).
Financial Rigor: Experience with systems where code correctness has direct financial impact—metering, ledgers, or payment gateways.
Operational Instinct: A track record of technical leadership through influence, driving architecture for multi-team platforms from design to production.
Experience with systems where correctness has financial consequences — metering, ledgers, billing, reservations, or comparable transactional domains
Strong experience with relational and NoSQL data models (especially with complex transactional data), plus working knowledge of streaming and analytical stores for time-windowed aggregation
Experience designing policy- or rules-driven systems that change through configuration rather than deploys — versioning, safe rollout, and rollback of behavioural change
Sound judgement on the enforcement trade-off: how to set thresholds and actions that stop real harm without punishing legitimate customers, and how to measure whether you got it right
Familiarity with monitoring, SLOs, and root-cause analysis in production environments
Familiarity with hallucination mitigation strategies (e.g., code constraints, test scaffolding, embeddings, context injection)
Track record of technical leadership without formal authority — driving cross-team architecture to a decision, writing designs others build from, and lifting the engineering standard around you
Excellent communication and cross-functional collaboration skills, including with Finance, Risk and Support stakeholders
A pragmatic, hands-on engineer who balances long-term vision with iterative execution
Experience with ElasticSearch cluster management, multi-tenant indexing, or lifecycle policies
Familiarity with ClickHouse or other OLAP systems for analytics
Background in event-driven systems, data pipelines, or message brokers
Contributions to open-source or public technical writing on database performance or systems design
Abuse & Fraud Expertise: Experience building rules engines, velocity checks, or adversarial defense systems at massive scale.
Modern Data Stack: Hands-on with high-throughput tools like Go, Redis (Lua), ClickHouse, and Flink for real-time aggregation.
Hyper-Growth Scaling: Experience navigating the transition from early architecture to load-bearing platform at scale.
Exposure to stream processing at scale (Flink, Kafka Streams, Beam) and columnar analytical stores (ClickHouse, BigQuery) for real-time aggregation
Hands-on experience with Node.js and Go in production services, particularly on latency-sensitive or high-throughput paths
Experience working with AI Tools, Billing & Metered Systems
Contributions to open-source tools, internal platforms, or engineering blogs
EEO Statement:
The company is an Equal Opportunity Employer. As an employer subject to affirmative action regulations, we invite you to voluntarily provide the following demographic information. This information is used solely for compliance with government recordkeeping, reporting, and other legal requirements. Providing this information is voluntary and refusal to do so will not affect your application status. This data will be kept separate from your application and will not be used in the hiring decision.
#LI-Remote #LI-NJ1