PySpark Developer
Experience: 3-7 Years
Location: PAN India
Key Skills
- Strong experience in PySpark, Apache Spark, Python, and SQL.
- Expertise in building and optimizing ETL/Data Processing Pipelines.
- Good understanding of Spark Architecture (RDDs, DataFrames, Spark SQL, DAGs).
- Experience with Hadoop Ecosystem (Hive, HDFS) and large-scale data processing.
- Exposure to AWS/Azure/GCP, object storage, and cloud-based data solutions.
- Knowledge of Data Modeling, Data Engineering, and Pipeline Design.
- Hands-on experience with Git and Linux/Unix environments.
- Strong debugging, performance tuning, and problem-solving skills.
Responsibilities
- Design, develop, and maintain scalable PySpark-based data pipelines.
- Build ETL workflows for ingesting, transforming, and loading large datasets.
- Optimize Spark jobs for performance, scalability, and reliability.
- Process data from multiple sources, including databases, APIs, HDFS, and cloud storage.
- Ensure data quality, validation, and governance standards.
- Troubleshoot and tune Spark applications for maximum efficiency.
- Collaborate with data engineers, analysts, and data scientists to deliver business solutions.
- Support production deployments and ongoing application maintenance.
- Follow security and best practices for enterprise data engineering.
Good to Have
- Spark Streaming / Structured Streaming
- Kafka or other messaging platforms
- Databricks, Airflow, or Azure Data Factory
- Data Warehousing concepts
- Agile/Scrum experience
Preferred: Candidates with strong hands-on PySpark development experience, cloud exposure, and expertise in building high-performance big data solutions.