Magnitudeminds AI Implementation

Careers/Engineering

Senior Data Engineer, ML/AI

Own the data layer that feeds every model we train and every agent we serve. Build ingestion and transformation across Snowflake, Databricks, Dagster, Airbyte, and keep training/serving data in sync.

Hybrid · Ho Chi Minh City, VietnamFull Time5+ years

About the role

We are hiring a Senior Data Engineer focused on the ML/AI data layer. This is distinct from our existing Data Platform Engineer role: you are specifically the person who makes sure our models train on good data and serve from the same features they trained on. You will design training-data pipelines, operate feature stores, manage Airbyte ingestion across dozens of sources, model training and serving tables in Snowflake or Databricks, and orchestrate the whole graph with Dagster. Your scorecard is data freshness, training-serving parity, and the speed at which an AI engineer can go from idea to a clean dataset.

Responsibilities

ML Data Pipelines

  • Design and operate training-data pipelines on top of Dagster: ingestion, cleaning, labeling, splits, snapshots.
  • Build feature engineering pipelines that land in a feature store (Feast or equivalent) with offline/online parity guarantees.
  • Partner with AI Engineers on dataset versioning, lineage and reproducibility for every trained model.
  • Implement synthetic-data, weak-supervision and active-learning loops where appropriate.

Cloud Data Platform

  • Own tables and transformations in Snowflake, Databricks (Delta Lake / Unity Catalog) or BigQuery for ML workloads.
  • Design bronze / silver / gold layers for training features, labels and evaluation datasets.
  • Maintain optimisation patterns: partitioning, clustering, OPTIMIZE / ZORDER, incremental refresh.
  • Manage cost and governance: query profiles, warehouse sizing, Unity Catalog / RBAC policies.

Ingestion & Integration

  • Operate Airbyte connector topology for SaaS, database and event ingestion; build custom connectors where needed.
  • Design CDC, incremental and full-refresh strategies with clear failure-recovery semantics.
  • Stand up schemas, contracts and quality gates at the ingestion boundary with asset checks in Dagster.
  • Integrate event streams (Kafka, Pub/Sub, NATS) into the batch / streaming-hybrid graph where low latency matters.

Quality, Observability & Collaboration

  • Define data-quality SLOs per asset; monitor freshness, null rate, distribution shift and schema drift.
  • Build automated eval-data pipelines that AI engineers depend on for regression testing models and agents.
  • Collaborate with AI Full-Stack and AI Engineer teams on retrieval and RAG dataset design.
  • Document pipelines, data contracts and domain semantics; level up the team on modern data engineering practices.

Qualifications

Must-Have Technical Expertise

  • 5+ years of data engineering, with 2+ years directly supporting ML or AI workloads.
  • Expert Python and advanced SQL; proficiency with dbt for modelling.
  • Hands-on experience with Dagster (asset graphs, sensors, schedules, asset checks) in production.
  • Direct production experience with at least one of: Snowflake, Databricks, BigQuery.
  • Direct production experience with Airbyte (or a comparable connector platform like Fivetran).
  • Understanding of feature-store patterns and training-serving parity.
  • Proficiency with AI-assisted development tools (Cursor, Claude Code, GitHub Copilot, or similar).

ML & Integration Skills

  • Exposure to at least one feature-store implementation (Feast, Tecton, or in-house).
  • Familiarity with ML lifecycle: experiment tracking, model evaluation, rollout.
  • Strong REST API design intuition for data access endpoints.
  • Comfort partnering with AI engineers on RAG / retrieval pipelines (chunking, embeddings, vector stores).

Preferred/Bonus

  • Experience building or contributing Airbyte connectors.
  • Deep work with Unity Catalog, MLflow on Databricks, or Snowpark.
  • Experience with streaming frameworks (Flink, Spark Structured Streaming, Beam).
  • Experience with data-contract tooling (Schemata, dbt-contracts, Great Expectations).
  • Strong Vietnamese and English communication skills.

Benefits

  • Competitive salary and performance incentives
  • Own the ML data layer end-to-end, from ingestion to training to evaluation
  • Work across modern cloud data platforms (Snowflake, Databricks)
  • Flexible work arrangements
  • A collaborative, innovative engineering team environment

Apply

Apply for Senior Data Engineer, ML/AI.

Send a CV and a short note on what you have shipped. We read every application and reply to all of them.

Apply by email