…

Associate III - Data Engineering

UST · IT Services & Consulting

  • Bangalore
  • On-site
  • Posted today
  • Data & AI

About the job

Data Pipeline Development: Design, build, and optimize robust, scalable, and efficient ETL/ELT data pipelines using Python and PySpark, primarily within Azure Databricks and Azure Data Factory. • Data Ingestion & Processing: Develop and manage processes for ingesting data from various sources (e.g., transactional databases, APIs, streaming sources) and transforming it into clean, usable formats for downstream consumption. • Data Quality & Monitoring: Implement comprehensive unit and integration test coverage for data pipelines. Establish and maintain monitoring, ing, and dashboarding solutions (e.g., Grafana) for data quality, pipeline health, and performance. • Cloud Infrastructure Management (OpenShift/Azure): Contribute to the setup, configuration, and maintenance of data-related infrastructure on OpenShift, ensuring deployment readiness and leveraging tools like HELM for application packaging and deployment. • CI/CD & Automation: Drive CI/CD best practices using GitHub Actions, ensuring automated testing (unit tests), build, and deployment processes for data solutions to environments like OpenShift. • SQL & Data Modeling: Develop and optimize complex SQL queries for data extraction, transformation, and loading. Apply strong data modeling principles for efficient data storage and retrieval in SQL Server and other data stores. • Azure Ecosystem Leverage: Utilize a broad range of Azure data and analytics services, including Azure Data Factory, Azure Databricks, Azure SQL Server, Azure Key Vault, and others to build comprehensive data solutions. • Performance Optimization: Proactively identify and resolve performance bottlenecks in data pipelines and databases through query optimization, indexing strategies, and efficient data processing techniques. • Collaboration & Documentation: Work closely with data scientists, analysts, and other engineering teams to understand data requirements. Create clear and concise documentation for data pipelines, architecture, and processes.