About the Role
The company is seeking an experienced Data Engineer with strong expertise in Apache Spark, Databricks, Scala, PySpark, Python, and SQL.
The engineer will design and maintain large-scale data pipelines across on-premises and cloud environments. The role also involves modernizing legacy workflows, supporting hybrid data architectures, integrating multiple storage systems, and improving the reliability and performance of production data processes.
Responsibilities
- Design, develop, and maintain scalable data pipelines using Apache Spark, Databricks, Scala, and PySpark.
- Build and support complex on-premises data workflows.
- Develop hybrid on-premises and cloud data integration solutions.
- Integrate data across HDFS, NAS, on-premises file shares, Amazon S3, and other storage platforms.
- Work with JSON, Parquet, CSV, Avro, fixed-length, and Excel data formats.
- Develop optimized SQL queries for data extraction, transformation, and loading.
- Connect to relational and non-relational databases.
- Implement efficient data extraction strategies.
- Develop and maintain workflow orchestration using Apache Airflow or similar tools.
- Write production-grade Python code for data processing, automation, and engineering utilities.
- Develop unit tests, integration tests, and data-quality tests for data pipelines.
- Troubleshoot pipeline failures and performance bottlenecks.
- Investigate data quality issues.
- Resolve complex integration problems across multiple systems.
- Support the migration and modernization of legacy on-premises processes.
- Contribute to hybrid and cloud data initiatives.
- Collaborate with Data Scientists, Analysts, Application Engineers, and other stakeholders.
- Prepare technical documentation covering pipelines, data flows, architecture, and data lineage.
- Support cloud integration initiatives, particularly within Azure environments.
- Use AI coding assistants and AI agents to improve development productivity and automate engineering tasks.
Requirements
- Strong hands-on experience with Apache Spark.
- Strong experience with Databricks.
- Strong Scala and Spark Scala development skills.
- Strong PySpark experience.
- Strong Python programming skills.
- Advanced SQL skills, including complex joins, query optimization, and performance tuning.
- Hands-on experience with Amazon S3.
- Experience working with HDFS.
- Experience with NAS and on-premises file systems.
- Experience integrating cloud and on-premises storage environments.
- Experience handling JSON, Parquet, CSV, Avro, fixed-length, and Excel files.
- Experience extracting data efficiently from multiple databases.
- Strong understanding of complex on-premises data workflows.
- Experience with multi-system integrations.
- Experience building hybrid on-premises and cloud data pipelines.
- Strong troubleshooting and production support skills.
Additional Skills
- Experience with Azure cloud services.
- Experience with Apache Airflow or similar workflow orchestration tools.
- Experience with automated unit and integration testing.
- Knowledge of data-quality validation and monitoring.
- Experience preparing data lineage and technical documentation.
Preferred Qualifications
- Working knowledge of Java.
- Experience with Prefect.
- Familiarity with React for internal tools or dashboards.
- Experience using AI coding assistants and AI agents.
- Experience in Pharmacy Benefit Management, healthcare, or a related domain.
