Data Engineer
EXL · IT Services & Consulting
- Gurugram, Haryana, India
- On-site
- Posted today
- Data & AI
About the job
Build and operate the data pipelines that feed the Entity Hub. This role lands all six in-scope sources into Fabric, implements standardization and transformation logic, and maintains the data quality checks and monitoring that the entity resolution engine depends on. Reliable, observable ingestion is the foundation the entire programmed rests on.
Documentation — produce and maintain source-to-target mappings, transformation logic documentation and lineage records
Skill Area
Specific Requirements
Core Engineering
Python, PySpark, advanced SQL, Delta Lake, distributed data processing
Microsoft Fabric
Data Factory pipelines and Copy Activity, Lakehouse, OneLake, Spark notebooks, Environments, Mirroring, Shortcuts
Data Integration
Batch and incremental ingestion, CDC patterns, watermarking, reprocessing strategies, schema-on-read for varied formats
Data Quality
Validation rule implementation, completeness/accuracy checks, alerting, exception workflows, reconciliation
Modelling
Bronze/Silver/Gold medallion layering, cleansing and conformance, standardization of names, addresses, dates and codes
Ops & Governance
Pipeline monitoring, lineage and metadata capture, access controls, technical documentation
Must-Have Qualifications
Nice-to-Have
Key Deliverables Owned
Dual Role / Complementary Skills
Complementary with the Entity Resolution engineering workstream — both are PySpark-on-Fabric disciplines, so this role can cross-train on Splink tuning and candidate-pair generation to provide cover. Also supports the Sr. Data Engineer (Lead) on identifier-spine construction, and can assist the VectorDB Engineer with document/attribute preparation in Phase 2.