Back to Job Portal
I

Python-Pyspark Developer

Infosys · HYDERABAD
Posted June 27, 2026 via Infosys

We are looking for an experienced Python PySpark Developer to design, develop, and optimize large-scale data processing systems. The ideal candidate will work on big data platforms, build scalable ETL pipelines, and process high-volume datasets using Spark and Python.

Key Responsibilities

Data Engineering & Development

Develop and maintain data pipelines using Python and PySpark

Process and transform large datasets in distributed environments

Build scalable ETL/ELT workflows

Big Data Processing

Work with Apache Spark (PySpark) for batch and real-time processing

Optimize Spark jobs for performance and efficiency

Handle structured and unstructured data

Data Integration

Ingest data from multiple sources:

Databases (SQL/NoSQL)

APIs

Files (CSV, JSON, Parquet)

Integrate with data platforms like:

Hadoop (HDFS)

Cloud (AWS, Azure, GCP)

Performance Optimization

Tune Spark jobs (partitioning, caching, parallelism)

Optimize SQL queries and transformations

Improve data processing efficiency and cost

Collaboration & Support

Work with data engineers, data scientists, and analysts

Translate business requirements into technical solutions

Participate in code reviews and agile development practices

Monitoring & Troubleshooting

Debug and resolve issues in data pipelines

Monitor job execution and data quality

Ensure reliability and availability of data workflows

• Primary skills:Python, Pyspark
Applying to this role?

Book a mock interview matched to your skills and get a written scorecard before the real thing.

Book interview prep