Technical SME Data Engineer
Quick Summary
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Technical SME Data Engineer based in the United States.
This role focuses on transforming large-scale regulatory data into actionable insights through advanced data engineering, statistical modeling, and machine learning. You will design and implement decision analysis and ML solutions within a cloud-native environment supporting federal regulatory initiatives. The position requires deep expertise in Apache Spark and large-scale data processing, along with the ability to translate organizational objectives into practical analytical solutions. You will also help ensure AI and ML work aligns with federal governance, documentation, and bias-testing requirements. Collaboration with program, policy, and technical teams will be central to communicating findings and model outcomes. This is a full-time opportunity based in McLean, Virginia, with potential remote flexibility depending on project requirements
The Technical SME Data Engineer will combine advanced data engineering and machine learning expertise to develop scalable analytical solutions while supporting regulatory, governance, and compliance objectives.
- Design, build, and implement decision analysis and machine learning models, including Decision Trees, Random Forests, Gradient Boosted Trees, Linear Regression, Collaborative Filtering, and K-Means.
- Develop large-scale data processing pipelines using Apache Spark core APIs, SparkSQL, and streaming capabilities.
- Perform data cleaning, transformation, validation, and comparative trend analysis across large datasets.
- Translate organizational goals and business requirements into practical machine learning models, statistical models, and pattern recognition solutions.
- Implement analytical solutions within cloud-native environments and support scalable data processing requirements.
- Ensure machine learning and AI initiatives comply with applicable federal AI governance policies, including documentation and bias-testing requirements.
- Contribute to open-source and community-driven solutions while maintaining source code through GitHub.
- Collaborate with program, policy, and technical stakeholders to communicate analytical findings, model results, and technical recommendations.
- Support data-driven initiatives that strengthen regulatory analysis and decision-making.
- Contribute subject matter expertise to technical discussions, solution development, and project delivery activities.
Requirements
~2 min readThe ideal candidate combines strong data engineering and machine learning experience with practical expertise in cloud-native technologies, large-scale analytics, and federal governance environments.
- 5+ years of professional experience working with decision analysis and machine learning algorithms in cloud-native environments.
- Strong hands-on experience with algorithms such as Decision Trees, Random Forests, Gradient Boosted Trees, Linear Regression, Collaborative Filtering, and K-Means.
- Strong understanding of Apache Spark architecture and internals, including core APIs, SparkSQL, high-level data access tools, and Spark streaming.
- Experience performing data cleaning and developing comparative trend analyses using large-scale datasets.
- Proven ability to translate organizational goals into working machine learning models, statistical models, or pattern recognition solutions.
- Experience working with open-source and community solutions and managing source code using GitHub.
- 5+ years of experience using Amazon Web Services EMR Spark is preferred.
- Experience with Zeppelin or similar data science interpreters is preferred.
- Familiarity with federal AI governance requirements, including model documentation and bias testing.
- Prior experience supporting federal financial regulators such as CFPB, FDIC, OCC, SEC, FRB, or NCUA is preferred.
- Experience working in a federal FISMA Moderate or comparable security and compliance environment is beneficial.
- Bachelor’s degree in Mathematics, Data Science, or a similar field required; a Master’s degree in Mathematics, Data Science, or a related discipline is strongly preferred.
- Strong analytical, problem-solving, communication, and collaboration skills.
- Ability to communicate technical findings clearly to program, policy, and technical stakeholders.
- Ability to work effectively in environments requiring security, regulatory, and governance awareness.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- October 2, 2026
- First seen
- October 2, 2026
- Last seen
- October 2, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- October 2, 2026
Signal breakdown
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.