Skip to main content

Rodrigo Monteiro Pereira · Data Platform & AI Lead Engineer

I build and govern data platforms; ML and agents go to production on top of them.

Data engineering and data science first; AI on top — measured, gated in CI, and every number on this page links to the artifact it came from.

About

Rodrigo Monteiro Pereira, Data Platform & AI Lead Engineer

I lead a Data & AI function end to end — architecture through execution. I started in C++ systems work in 2018, moved into data in 2021 and have led data teams since 2022, first in the financial market and today in telecoms, across on-prem and cloud.

My focus is the load-bearing part: medallion architecture, data quality, lineage, cost and access. AI and agents sit on top of that platform — with real evaluation suites, guardrails and CI regression gates, not demos.

Current focus

  • Governed lakehouse platforms — medallion, contracts, lineage
  • RAG and agents with evaluation suites and CI regression gates
  • MLOps — drift monitoring, leakage control, gated retraining
  • Governance under real regulators — LGPD, ANATEL, FCC, ANACOM
  • Data quality, master data and cost — the unglamorous half

Portfolio projects

Six projects across platform, ML and applied AI. Statuses below are honest — nothing here is described as finished before it is.

  • Data platform · Lakehouse

    Status: Live

    Open Finance LakeHouse

    51 financial series, batch and streaming, on an entirely open stack.

    The foundation the rest of the portfolio is built on. Public Brazilian macro and market data lands as Delta on MinIO, is conformed by Spark into a star schema and served as DuckDB marts — every lane generated from one source registry. Adding a series is a YAML entry, not new code.

    • 51 registered series across 10 source handlers (BACEN, IPEA, IBGE, Tesouro, ANBIMA, B3, Yahoo Finance) — medallion bronze → silver → gold on Delta Lake over MinIO, ending in 10 BI-ready marts.
    • Airflow 3 Asset scheduling derived from the registry at parse time — 13 DAGs and 51 bronze assets, one task per series, so a failing API withholds exactly one asset and raises exactly one attributable alert.
    • Spark
    • Structured Streaming
    • Delta Lake
    • Airflow 3
    • MinIO
    • Polars
    • DuckDB
    • Pandera
    • OpenLineage
    • OpenMetadata
    • Kubernetes

    The Kubernetes/GitOps layer that deploys this stack lives in a separate repository that holds cluster credentials. It stays private permanently and is not linked from here.

  • Applied AI · Retrieval, evaluation & guardrails

    Status: In progress

    RAG over BACEN Copom minutes

    A retrieval defect, found by measurement and fixed by 120 lines of regex.

    RAG over 30 public minutes of the Copom, Brazil's monetary policy committee, built around an evaluation harness rather than a demo. Thirty documents each contain a near-identical decision paragraph, so dense retrieval confidently returned the right paragraph from the wrong meeting. Seven measured retrieval arms later, that defect is gone — and the component that fixed it was not the one anybody would have bet on.

    • The defect, measured: the dense baseline scored MRR 0.191 and got the right meeting at rank 1 on 4 of the 41 questions that name their meeting outright. It was not confused about the topic — it was answering a question nobody asked.
    • The fix, measured on the same question set: MRR 0.191 → 0.741, hit@5 0.959, rank-1 meeting correctness 0.098 → 1.000. It came from ~120 lines of regex that read the date out of the question and compile it into a Qdrant payload filter, at 41/41 hint precision and negative latency cost.
    • Python
    • Qdrant
    • bge-m3
    • BM25
    • cross-encoder
    • Ollama
    • Presidio
    • Langfuse
    • FastAPI
    • DuckDB

    Every number above comes from a committed JSON report in the repo. The gold set behind them is 56 draft pairs, not human-validated — so read the movement between arms, not the absolute level.

  • Machine learning · MLOps

    Status: Live

    Hourly energy demand forecasting & drift

    The forecast is the easy half; the loop that decides when to retrain is the project.

    Hourly electricity demand forecasting for PJM, measured on two years of real load from the EIA API and wrapped in the loop that actually matters in production: a scheduled job runs every day, unattended — re-scoring the frozen champion against the actuals that arrived, computing four kinds of drift, deciding whether to retrain, and publishing the numbers back to the repository by pull request. Drift accumulates in public, week after week, and can be pointed at.

    • The forecast, measured on real data: 2.72% MAPE against 11.73% for a seasonal-naive baseline on identical folds — 55 walk-forward refits over eight weeks, winning 24 of 24 horizons. The panel behind it is two years of hourly PJM demand with no missing hours, and every artifact derived from it carries "is_real": true.
    • The ablation that pays for a design choice: removing the archived day-ahead weather forecasts and refitting costs 41% of the improvement over the baseline. That measurement — not an argument — is what justifies the publication-masked forecast features.
    • Python
    • LightGBM
    • MLflow
    • FastAPI
    • pandas
    • Parquet
    • Evidently
    • React
    • ECharts
    • GitHub Actions

    The evidence base is stated at its real size: eight summer weeks of real PJM load, and no winter peak has been forecast yet.

  • Machine learning · Research benchmark

    Status: Complete

    Vision Transformers for vertebra landmarks

    169 training runs collapsed into one grid, every number re-derivable from its archive.

    A controlled benchmark of Vision Transformers for localising two cervical-spine landmarks in videofluoroscopy frames: nine models by six configurations, 48 of 54 cells filled, a 95% CI on every cell. The dataset is licensed medical imaging and is not in the repository — and the entire analysis verifies without it: no GPU, no data, one command.

    • 169 archived training runs collapsed into the grid — with a single traceable run hash behind every number, and a checker that re-derives every published figure from the archive or fails: `aggregate_results.py --check`.
    • The finding: freezing the backbone is the most damaging single choice measured, and the best fine-tuned cell reaches 5.18 ± 0.40 px against the 6.31 px baseline.
    • Python
    • PyTorch
    • Hugging Face Transformers
    • NumPy
    • Matplotlib
    • GitHub Actions

    The underlying dataset is licensed medical imaging (CC BY-NC-SA) and is not in the repository. The code, the run archive, the figures and the tests are — nothing here reconstructs the data, and CI enforces that.

  • Applied AI · End-to-end analytics product

    Status: Complete

    Shopping Intelligence

    Every metric, forecast and agent answer traces to deterministic data and services.

    An end-to-end analytics platform over a monthly mall panel, built for a technical challenge and delivered whole: a medallion pipeline with an ODCS contract, Pandera validation and SHA-256 provenance; reproducible EDA and PCA; four model candidates under identical rolling-origin folds; and the same executive view served twice — a Next.js dashboard and a versioned Power BI project — from the same gold layer.

    • Model selection that resists its own marketing: four candidates (Dummy, Bayesian Ridge, GAM, CatBoost) on identical expanding temporal folds; the published champion is a regularised GAM chosen by a parsimony rule — and the repo states that its advantage is not statistically separable, publishing the paired-bootstrap 95% CI of the RMSE difference, [-69; +15] against CatBoost, in the champion metadata.
    • The agent never computes a metric. A LangGraph graph scopes the question, picks among allowlisted, validated tools, executes, composes and validates — the LLM writes from the returned payload, with numeric grounding checked, a behavioural evaluation suite and a multi-LLM benchmark in the repo.
    • Python
    • pandas
    • scikit-learn
    • CatBoost
    • Pandera
    • DuckDB
    • FastAPI
    • Next.js
    • LangGraph
    • Power BI
    • Docker

    The panel is one annual cycle of a public 120-record base, so no long-horizon seasonal claim is made; month and event are perfectly confounded, and every figure here is a validation measurement, not a causal effect.

  • Master data management

    Status: Planned

    MDM / entity resolution

    One customer, five systems, no agreement — resolved into a golden record.

    Entity resolution over dirty, overlapping public datasets: blocking to make the comparison tractable, then ML and embedding based matching, ending in a single governed golden record. Designed to run on top of the Open Finance LakeHouse rather than as a separate stack. Not started — this card is a plan, and it says so.

    • Blocking strategy first, so the pairwise comparison stays tractable at scale.
    • Matching by ML and embeddings rather than hand-tuned string rules alone.
    • Python
    • Spark
    • Embeddings
    • Delta Lake
    • DuckDB

    Not started — no code yet

Experience

  1. Data Platform & AI Lead Engineer

    Feb 2026 – Present
    TelecallRio de Janeiro, Brazil

    Owns the Data & AI function end to end — architecture through execution — reporting directly to the IT Director, across an analytics estate spanning terabytes and three regulatory jurisdictions.

    • Databricks
    • Spark
    • Data governance
    • AI agents
    • GitOps
  2. Data & AI Consultant

    Aug 2023 – Feb 2026
    Consultoria IndividualRio de Janeiro, Brazil

    Independent data and AI consulting, between the Faros and Telecall roles.

    • Data Team Lead

      Jul 2022 – Aug 2023
      Faros Private InvestimentosRio de Janeiro, Brazil

      Led the data team, designing and shipping scalable data infrastructure for a private investment firm.

      • Databricks
      • Airflow
      • Spark
      • LakeFS
      • Python
    • Data Scientist

      Aug 2021 – Jul 2022
      Faros Private InvestimentosRio de Janeiro, Brazil

      Built predictive models and analytics products to sharpen commercial strategy and improve the client experience.

      • Python
      • PostgreSQL
      • Power BI
      • React
      • Next.js
    • Software Developer (Intern)

      Aug 2018 – Aug 2020
      Instituto TecgrafRio de Janeiro, Brazil

      Built modelling and simulation software for oil reservoirs in C++, OpenGL and Qt.

      • C++
      • OpenGL
      • Qt
      • 3D modelling
      • Simulation

    Skills

    Platform & governance

    • Apache Spark
    • Delta Lake
    • Apache Airflow
    • Databricks
    • LakeFS
    • MinIO / S3
    • Medallion architecture
    • Pandera
    • Great Expectations
    • OpenLineage
    • OpenMetadata

    Data science & ML

    • Python
    • pandas
    • Polars
    • scikit-learn
    • MLflow
    • Statistical modelling
    • Time-series validation
    • Drift monitoring (PSI / KS)

    AI & agents

    • RAG
    • Hybrid retrieval
    • Reranking
    • Evaluation suites
    • LLM-as-judge
    • Guardrails
    • Tracing (Langfuse)
    • Text-to-SQL
    • MCP

    Query & storage

    • SQL
    • DuckDB
    • PostgreSQL
    • Parquet
    • Power BI

    Cloud & DevOps

    • Azure
    • AWS
    • Docker
    • Kubernetes
    • FluxCD
    • Talos
    • Terraform
    • GitHub Actions
    • CI/CD

    Languages & interfaces

    • Python
    • C++
    • TypeScript
    • React
    • Next.js
    • Tailwind CSS
    • Recharts

    Background

    MSc in Data Science and Artificial Intelligence

    Pontifícia Universidade Católica do Rio de Janeiro (PUC-Rio)

    Aug 2025 – in progress

    BSc in Computer Science

    Pontifícia Universidade Católica do Rio de Janeiro (PUC-Rio)

    2016 – 2021

    Spoken languages

    • Portuguese — native
    • English — fluent
    • Spanish — fluent

    Get in touch

    Open to conversations about data platform, governance and applied AI roles.