Data Engineer Senior Consultant
NTT DATA · IT Services & Consulting
- Bangalore, IN-KA, India
- On-site
- Posted today
- Data & AI
About the job
Build scalable data ingestion pipelines within Databricks using Python, PySpark to process and extract data from PDF files
Implement layout analysis tools (Unstructured.io, LlamaIndex, pdfplumber) to accurately extract text, titles, and embedded tables from complex multi-column documents.
Develop high-accuracy Named Entity Recognition (NER) layers leveraging a mix of open-source frameworks (spaCy, Hugging Face Transformers like LayoutLM) and cloud foundation models (AWS Bedrock).
Design, partition, and maintain optimal target schemas within Databricks Delta Lake tables to support downstream analytics, reporting, and vector search indexing.
Required Skills
Apache Spark/PySpark engineering on Databricks;
Delta Lake architecture, schema enforcement/evolution, partitioning and performance optimization;
Python-based ETL/ELT and scalable batch ingestion;
PDF parsing and layout-aware extraction using Unstructured.io, LlamaIndex and pdfplumber;
NLP/NER pipelines with spaCy, Hugging Face Transformers and LayoutLM;
AWS Bedrock foundation-model integration;