…

Python - Data Processing

Infosys · IT Services & Consulting

  • Bangalore, India
  • On-site
  • Posted 10 days ago
  • Software Engineering

About the job:

Step into a role where your Python expertise directly powers reliable, high-quality data movement across systems. As part of a collaborative engineering team, you’ll help design and build scalable data processing solutions that enable faster insights and smoother operations. You’ll work closely with stakeholders to understand data needs, translate them into robust pipelines, and continuously improve performance and stability. This is a great opportunity for someone who enjoys solving real-world data challenges, values clean and maintainable code, and takes pride in delivering dependable systems. If you’re motivated by ownership, teamwork, and building data workflows that others can trust, you’ll feel right at home here.

Responsibilities

Key Responsibilities:

Build and maintain Python-based data processing components for structured and semi-structured datasets.
Develop and support ETL workflows to ingest, transform, validate, and load data into target systems.
Design and optimize SQL queries for data extraction, transformation, reconciliation, and reporting needs.
Implement reliable data pipelines with proper logging, error handling, retries, and monitoring hooks.
Perform data quality checks, anomaly detection rules, and reconciliation to ensure accuracy and completeness.
Troubleshoot pipeline failures and performance bottlenecks; drive root-cause analysis and permanent fixes.
Collaborate with cross-functional teams to gather requirements and deliver incremental improvements.
Maintain clear technical documentation for pipeline design, data mappings, and operational runbooks.

Technical requirements

Primary skills: Python - Data Processing/Technology->Big Data - Data Processing->PySpark,Technology->OpenSystem->Python - OpenSystem->Python

Additional responsibilities

Minimum Qualifications:

Bachelor’s degree (or equivalent) in Engineering/Computer Science/IT or related field (BTech/BE/BSc or equivalent).
2–3 years of hands-on experience in Python for data processing and automation.
Practical experience building ETL processes and working with data pipelines end-to-end.
Strong SQL skills including joins, aggregations, subqueries, and performance-aware query writing.
Ability to write clean, maintainable code and follow basic engineering practices (version control, reviews, testing mindset).

Preferred Qualifications:

Experience designing scalable pipeline patterns (incremental loads, CDC concepts, partitioning, backfills).
Familiarity with Python data libraries and processing approaches (e.g., Pandas, batch processing patterns).
Exposure to orchestration/scheduling concepts and operationalizing pipelines for reliability and observability.
Experience working with large datasets and optimizing end-to-end pipeline performance (I/O, SQL tuning, compute efficiency).
Proven ability to collaborate with stakeholders, translate requirements into technical solutions, and deliver within timelines.

Good to have skills:

Pandas, NumPy, Apache Airflow, Spark (PySpark), Linux/Shell Scripting

Educational requirements

MCA,MSc,MTech,Bachelor of Engineering,BTech

Preferred skills

Technology->OpenSystem->Python - OpenSystem->Python,Technology->Big Data - Data Processing->PySpark

Experience: 2-3 years