Engenheiro de Dados Pleno
Quick Summary
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Engenheiro de Dados Pleno based in Brazil.
This role focuses on building reliable, scalable data foundations for modern AI and analytics solutions. You will design and operate ETL/ELT pipelines using Databricks, Delta Lake, and lakehouse architectures. The position combines structured and unstructured data engineering, including the ingestion and preparation of PDFs, images, and text for RAG and evaluation workflows. You will contribute to data curation, anonymization, vector indexing, embeddings, and versioned datasets. The role also involves connecting data platforms with inference services and APIs while balancing performance, cost, monitoring, and reliability. Automation, data quality, CI/CD, and continuous improvement will be central to your work. You will join a dynamic, technology-focused environment where modern cloud and AI practices are used to solve complex business challenges.
The role is responsible for designing and operating data pipelines and platforms that support AI-enabled applications, ensuring data quality, governance, scalability, and operational reliability.
- Design and build ETL/ELT pipelines using Databricks or open-source technologies, ingesting and transforming historical data through a bronze, silver, and gold medallion architecture on Delta Lake.
- Ingest, standardize, and enrich unstructured documents such as PDFs and images with appropriate metadata.
- Curate and anonymize data in accordance with LGPD requirements, producing clean datasets for RAG, few-shot workflows, and testing.
- Build and operate vector indexing pipelines, including chunking, embedding generation, incremental index updates, and version control.
- Prepare versioned datasets and golden sets for evaluation frameworks and accuracy benchmarking.
- Structure persistence for feedback cycles, including positive and negative feedback, justifications, and resolution status, supporting quality and SLA dashboards.
- Integrate data platforms with inference services, APIs, and other systems while considering performance, cost, monitoring, and alerting.
- Automate data workflows through Jobs and Workflows, CI/CD pipelines, and data quality testing.
- Support continuous improvement of data engineering practices, reliability, and operational efficiency.
Requirements
~1 min readThe ideal candidate combines strong data engineering fundamentals with practical Databricks and cloud experience, plus an understanding of modern data and AI workflows.
- Advanced Python and SQL skills.
- Solid knowledge of data modeling and medallion/lakehouse architecture.
- Experience building data pipelines for unstructured data, including PDFs, images, and text, and preparing these datasets for RAG applications.
- Familiarity with vector databases and embeddings, such as Databricks Vector Search, pgvector, or similar technologies.
- Experience with data quality, data testing, and versioning practices.
- Hands-on experience with Databricks, including Delta Lake, Workflows/Jobs, notebooks, PySpark, and Spark SQL.
- Experience with Git and CI/CD practices.
- Professional experience working in cloud environments such as Azure, AWS, or GCP.
- Knowledge of LGPD and the appropriate handling of sensitive data.
- Experience with Unity Catalog in governed environments is desirable.
- Familiarity with MLflow, Databricks Model Serving, or Mosaic AI is a plus.
- Experience with Delta Live Tables, Lakeflow, or Auto Loader is desirable.
- Experience in insurance or financial services environments is an advantage.
- Familiarity with Terraform or Databricks Asset Bundles is a plus.
- Databricks Data Engineer Associate or Professional certification is desirable.
- Proactive approach, problem-solving mindset, adaptability, and willingness to continuously learn new technologies.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- October 2, 2026
- First seen
- October 2, 2026
- Last seen
- October 2, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- October 2, 2026
Signal breakdown
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.