Retour à la recherche
Y
YochanaSource d’offres vérifiée

Data Engineer - ETL, AI

Offre en anglais

Position Name – Data Engineer - AI Type of hiring – Subcon Location – Brampton, ON (Onsite) Job Description: The data backbone owner who ensures our AI systems have clean, structured, real-time data to reason over. You will design ingestion pipelines, vector indexing infrastructure, and governance layers that keep our RAG and agent memory systems accurate and compliant. About the Role You will architect ETL pipelines, manage vector databases, enforce governance, and build real-time data flows that continuously update embeddings and indexes. Your work ensures our AI agents operate with fres…

  • Sur place
  • ONTARIO
  • Publié 24 juill. 2026
  • Postuler avant le 23 août 2026
  • 1 poste

Résumé du poste

Position Name – Data Engineer - AI Type of hiring – Subcon Location – Brampton, ON (Onsite) Job Description: The data backbone owner who ensures our AI systems have clean, structured, real-time data to reason over. You will design ingestion pipelines, vector indexing infrastructure, and governance layers that keep our RAG and agent memory systems accurate and compliant. About the Role You will architect ETL pipelines, manage vector databases, enforce governance, and build real-time data flows that continuously update embeddings and indexes. Your work ensures our AI agents operate with fresh, trustworthy information. Data Pipelines: Build ETL flows for structured/unstructured data, ensuring normalization, deduplication, and semantic consistency. Vector Infrastructure: Manage pgvector, Azure AI Search, Redis vector indexing, and hybrid search layers. Data Governance: Implement zero-trust access, privacy controls, and compliance within AI context pipelines. Real-time Processing: Build event-driven architectures that continuously refresh embeddings and indexes. Required Qualifications Deep experience with distributed data systems, SQL, and orchestration tools. Experience tuning high-throughput database infrastructure. Knowledge of Google’s GECX is a plus. Familiarity with chunking strategies and embedding models. Skillset Requirements ETL & Data Modeling: Designing pipelines for structured/unstructured data, normalization, deduplication, and semantic consistency. Vector Databases: pgvector, Redis, Azure AI Search, hybrid search, and index optimization. Distributed Data Systems: Kafka, Spark, Flink, or similar event-driven architectures. Data Governance: Zero-trust access, privacy controls, compliance, and auditability. Real-time Embedding Updates: Event-driven refresh pipelines for RAG and agent memory systems. Chunking & Embeddings: Semantic chunking, metadata tagging, and embedding model selection. Search Infrastructure: BM25, hybrid search, inverted indexes, and ranking algorithms. Performance Tuning: High-throughput read/write optimization. Data Quality & Lineage: Validation, schema enforcement, and lineage tracking (e.g., Great Expectations, OpenLineage).

Ce que vous ferez

Architect and manage ETL pipelines and vector indexing infrastructure to support AI systems and RAG memory. Ensure data governance, real-time embedding updates, and the delivery of clean, structured data for AI agents.

Exigences

Requires deep experience with distributed data systems, SQL, and orchestration tools, specifically for high-throughput database infrastructure. Candidates should be familiar with chunking strategies, embedding models, and event-driven architectures.

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • ETL
  • Data Modeling
  • Vector Databases
  • Pgvector
  • Redis
  • Azure AI Search
  • Kafka
  • Spark
  • Flink
  • Data Governance
  • RAG
  • Semantic Chunking
  • Hybrid Search
  • Performance Tuning
  • Data Quality
  • SQL

Renseignements supplémentaires

Expérience minimale
5+ ans
Postuler avant le
23 août 2026