Senior Big Data Engineer – Hadoop / AWS EMR, Spark + Scala & Kafka
Virtues
Job description
Job Summary
We are seeking an experienced Senior Big Data Engineer with strong expertise in Hadoop, AWS EMR, Spark with Scala, Kafka, and large-scale data processing.
Candidates must have strong hands-on experience with either traditional Hadoop ecosystem platforms (such as Cloudera, Hortonworks, or MapR) OR AWS EMR (Amazon Elastic MapReduce). Experience with both is preferred but not required.
The ideal candidate will have extensive experience designing, developing, and supporting scalable batch, real-time, and streaming data pipelines across large-scale data platforms. Experience with Hadoop platform administration and GCP BigQuery is a plus.
Key Responsibilities
· Design, develop, implement, and maintain scalable data pipelines using Hadoop ecosystem technologies and/or AWS EMR.
· Develop and optimize Apache Spark applications using Scala for large-scale distributed data processing.
· Build data pipelines supporting batch, real-time/event-driven, and streaming processing.
· Build data ingestion solutions for structured, semi-structured, and unstructured data from multiple sources.
· Design and implement Kafka-centric event processing and real-time data pipelines.
· Develop streaming transformations using Spark Structured Streaming and/or Spark Streaming.
· Work with distributed storage technologies such as HDFS and/or Amazon S3.
· Perform data validation, cleansing, enrichment, transformation, and integration across raw, curated, and publishing layers.
· Work with relevant Hadoop/EMR technologies including Hive, HBase, YARN, MapReduce, Spark, Kafka, and related ecosystem components.
· Monitor, troubleshoot, and optimize Hadoop clusters and/or AWS EMR workloads, Spark applications, resource utilization, and processing performance.
· Collaborate with architects, application teams, business stakeholders, and platform teams to translate requirements into scalable technical solutions.
Required Qualifications
· 7+ years of Big Data/Data Engineering experience with large-scale distributed data platforms.
· Strong hands-on experience with either Hadoop ecosystem platforms OR AWS EMR. Experience with both is a plus, but not required.
· 5+ years of strong hands-on Apache Spark experience, including distributed data processing using Scala.
· Strong experience with Apache Kafka and event-driven data pipelines.
· Experience with Spark Structured Streaming and/or Spark Streaming.
· Strong understanding of batch processing, stream processing, and event-driven architecture.
· Strong knowledge of distributed data processing concepts and technologies such as HDFS/S3, Hive, YARN, HBase, and MapReduce.
· Experience designing and developing large-scale data ingestion, transformation, and integration pipelines.
· Experience with data modeling, data transformation, data quality, validation, and error handling.
· Strong understanding of Spark architecture, performance tuning, partitioning, memory management, and distributed processing.
· Strong analytical, debugging, troubleshooting, and root-cause analysis skills.
· Excellent communication and collaboration skills.
Preferred / Nice-to-Have
· Experience with both on-premises Hadoop and AWS EMR environments.
· Experience with Cloudera, Hortonworks, or MapR.
· Experience with Hadoop platform administration, cluster configuration, monitoring, troubleshooting, upgrades, and capacity planning.
· Experience with AWS services such as Amazon S3, AWS Glue, IAM, CloudWatch, Lambda, and Step Functions.
· Experience migrating or modernizing Hadoop/Cloudera workloads to AWS EMR.
· Experience with HiveQL, Pig, Sqoop, Oozie, ZooKeeper, Flume, and Hue.
· Experience with GCP BigQuery.
· Experience with workflow/orchestration tools such as Apache Airflow.
Key Skills
Hadoop OR AWS EMR | Apache Spark | Scala | Kafka | Spark Structured Streaming | HDFS / Amazon S3 | Hive | YARN | HBase | MapReduce | Batch Processing | Streaming | Distributed Data Processing
Pay: $100,000.00 - $110,000.00 per year
Work Location: In person