2mo ago
$176,400 – $242,550/yr

Reinforcement Learning Engineer

United StatesUnited StatesRemotemid
OtherEngineer
3 views0 saves0 applied

Quick Summary

Overview

Founded in 2012, Bugcrowd is the preemptive security platform that unifies exposure discovery and assessment, offensive testing,

Technical Tools
OtherEngineer

Founded in 2012, Bugcrowd is the preemptive security platform that unifies exposure discovery and assessment, offensive testing, and intelligence shaped by AI and human insight to help organizations avoid, discover, and validate real-world risk. Bugcrowd helps security teams move faster by identifying the exposures that matter most so they can act first and stay ahead of attackers. By combining the power of humans and AI, teams can preempt attack paths and prevent breaches. Based in San Francisco and New Hampshire, Bugcrowd is supported by General Catalyst, Rally Ventures, Costanoa Ventures, and others. Visit www.bugcrowd.com.

The Bugcrowd RL and Reasoning Team focuses on pushing the boundaries of autonomous cybersecurity by building authentic, verifiable reinforcement learning environments for world-leading foundational AI companies. As a Reinforcement Learning Engineer specializing in Reinforcement Learning from Verifiable Rewards (RLVR), you will design and scale automated verification pipelines that transform real-world software vulnerabilities into deterministic reward functions. In this role, you will bridge the gap between low-level security analysis and modern LLM reasoning models, engineering environments where AI agents learn to discover, exploit, and remediate software vulnerabilities with mathematical certainty. Instead of relying on subjective human feedback, your work directly powers the rigorous, verifiable reward signals that teach next-generation frontier AI models how to master complex cybersecurity domain logic. You will work at the intersection of fuzzing, dynamic program analysis, system exploitation, and scalable ML infrastructure to shape the safety and offensive/defensive capabilities of future artificial intelligence. 

Responsibilities

~1 min read
  • →Design, build, and deploy high-throughput RLVR (Reinforcement Learning from Verifiable Rewards) environments that evaluate LLM action sequences against deterministic execution outcomes.
  • →Develop automated test harnesses, sandboxes, and verification engines that convert complex vulnerability research (e.g., memory corruption, web security, logic bugs) into binary pass/fail reward signals.
  • →Integrate Bugcrowd’s Mayhem automated analysis platform and real-world vulnerability feeds into continuous, scalable RL environment generation pipelines.
  • →Architect safe, isolated, and highly reproducible execution environments (using Docker, BuildKit, or Nix) capable of running thousands of simultaneous agent-driven exploitation and patching trajectories.
  • →Collaborate directly with researchers at frontier AI labs including Anthropic, OpenAI, and Cohere to define standard benchmark formats, observation spaces, and verifiable evaluation metrics for cybersecurity tasks.
  • →Implement precise telemetry, ground-truth verification algorithms, and trajectory logging to analyze agent reasoning paths and prevent reward hacking or false positives.
  • →Build low-level instrumentation and debugging tools to monitor memory states, process executions, and network behaviors during agent interaction cycles.
  • →Optimize infrastructure performance and environment reset latency to support massive-scale parallel sampling and distributed RL training workflows.
  • →Benchmark and evaluate frontier AI model performance across diverse offensive and defensive security challenges, such as automated fuzzing, exploit payload generation, and patch validation.
  • Understanding of RL training workflows used by modern LLM systems, specifically execution-based feedback or Reinforcement Learning from Verifiable Rewards (RLVR).
  • Proficiency developing applications in Python and low-level systems programming in C, with Rust experience being a strong plus.
  • Solid understanding of software vulnerabilities, binary exploitation, fuzzing methodologies, or program analysis.
  • Experience with DevOps pipelines (e.g., GitHub Actions), reproducible builds (Docker, BuildKit, Nix), and comfort working with Linux systems and low-level debugging.
  • Experience working with or building benchmark environments (e.g., CTFs, SWE-bench, security challenges, or execution sandboxes).

Nice to Have

~1 min read
  • Experience designing custom reward functions, ground-truth verifiers, or automated grading engines for AI safety and reasoning models.
  • Background in low-level program analysis tools, sanitizers (e.g., ASan/MSan), compiler instrumentation, or automated exploit generation tools.
  • Proven track record of participating in or developing competitive cybersecurity benchmarks, CTFs, or open-source AI evaluation frameworks.

Requirements

~1 min read

The ideal candidate must be able to complete all physical requirements of the job with or without reasonable accommodation.

Sitting and / or standing - Must be able to remain in a stationary position 50% of the time

Carrying and / or lifting - Must be able to carry / move laptop as needed throughout the work day.

Environment - remote, work-from-home 100% of the time.

At Bugcrowd, we strive for fairness, equality and to create an environment that allows our people to perform at their very best. Our compensation philosophy is to foster a collaborative community that rewards, attracts and retains the best possible talent. The provided salary details are based on US national averages and we retain the flexibility to tailor to the needs of the business.

The national estimate for the current base range for the position of $176,400 - $242,550.

This position may also be eligible to participate in a discretionary bonus program or commission plan, subject to the rules governing the program, whereby an award, if any, depends on various factors, including, without limitation, individual and organizational performance.

  • At Bugcrowd, we understand that diversity in the workplace is vital to a company’s success and growth. We strive to make sure that people are included and have a sense of being part of making Bugcrowd not only a great product but a great place to work.
  • We regularly hear from both customers and researchers that Bugcrowd feels like a family, and we strive to maintain that internally as well.
  • Our team consists of a broad range of people: musicians, adventure sports junkies, nature lovers, parents, cereal enthusiasts, night owls, cyclists, artists—you get the point.

At Bugcrowd, we are solving security threats and vulnerabilities that are relevant to everyone, therefore we believe solving these problems takes all kinds of backgrounds. We value the perspectives and experiences people from underrepresented backgrounds bring.

This position has access to highly confidential, sensitive information relating to the technologies of Bugcrowd. It is essential that the applicant possess the requisite integrity to maintain the information in the strictest confidence.

The company is authorized to obtain background checks for employment purposes under state and federal law. Background checks will be conducted for positions that involve access to confidential or proprietary information (including trade secrets).

Background checks may include Social Security verification, prior employment verification, personal and professional references, educational verification, and criminal history. Applicants with conviction histories will not be excluded from consideration to the extent required by law.

Any personal data you submit in connection with your application will be processed in compliance with Bugcrowd's Privacy Policy, which you may review here: https://www.bugcrowd.com/privacy.

Bugcrowd is EOE, Disability/Age Employer. 

Individuals seeking employment at Bugcrowd are considered without regards to race, color, religion, national origin, age, sex, marital status, ancestry, physical or mental disability, veteran status, gender identity, or sexual orientation. 

Bugcrowd is committed to the full inclusion of all qualified individuals. In keeping with our commitment, Bugcrowd will take the steps to assure that people with disabilities are provided reasonable accommodations. Accordingly, if reasonable accommodation is required to fully participate in the job application or interview process, to perform the essential functions of the position, and/or to receive all other benefits and privileges of employment, please contact HR at ADA at bugcrowd.com.

Apply at: https://www.bugcrowd.com/about/careers/

Location & Eligibility

Where is the job
United States
Remote within one country
Who can apply
US

Listing Details

Posted
July 15, 2026
First seen
July 15, 2026
Last seen
October 10, 2026

Posting Health

Days active
86
Repost count
0
Trust Level
49%
Scored at
October 10, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust

Bugcrowd is the #1 crowdsourced security platform.

Employees
125
Founded
2012
View company profile
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Reinforcement Learning Engineer$176k–$243k