LigaData
LigaData

Junior Data Engineer

Junior Data Engineer - Jordan Office

Job Overview

As a Junior Data Engineer at Ligadata, you will support the development, maintenance, and monitoring of data pipelines and Big Data solutions. You will work with senior engineers and cross-functional teams to ensure reliable data processing, data quality, and timely delivery.

The role requires a good foundation in SQL, Linux, Shell scripting, data analysis, and Big Data technologies, with a strong willingness to learn and troubleshoot within a production data environment.

Responsibilities

  • Develop and maintain ETL/ELT data pipelines.
  • Write and optimize SQL queries for data processing, validation, and analysis.
  • Support data workflows using Apache Airflow.
  • Work with Big Data technologies such as Hadoop, HDFS, Hive, Spark, Presto/Trino, Kafka, and HBase.
  • Perform data validation, reconciliation, and data-quality checks.
  • Develop scripts and automation using Shell/Bash and Python.
  • Monitor data pipelines and assist in troubleshooting job failures and production issues.
  • Work with structured and semi-structured data formats such as Parquet, JSON, CSV, and Avro.
  • Support applications and data workloads running on Kubernetes (K8s).
  • Analyze data to identify inconsistencies, anomalies, and operational issues.
  • Participate in code reviews, documentation, and continuous improvement activities.
  • Collaborate with senior engineers, DevOps, QA, database, and analytics teams.
  • Use AI-assisted engineering tools to support development, troubleshooting, documentation, and data analysis while validating generated results before use.


Qualifications

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field.
  • 1–3 years of experience in Data Engineering, Big Data, Software Engineering, or a related role.
  • Good knowledge of SQL and relational database concepts.
  • Good understanding of Linux and Shell/Bash scripting.
  • Basic knowledge of the Hadoop ecosystem, including HDFS and Hive.
  • Familiarity with Spark, Presto/Trino, and Apache Airflow.
  • Basic understanding of Kafka and distributed data-processing concepts.
  • Familiarity with Kubernetes and containerized environments.
  • Knowledge of at least one programming language, preferably Python, Scala, or Java.
  • Basic understanding of ETL/ELT, data warehousing, data quality, and data analysis.
  • Familiarity with Git and software-development practices.
  • Knowledge and practical experience using AI tools such as ChatGPT, GitHub Copilot, or similar tools for engineering tasks.
  • Good analytical, troubleshooting, and problem-solving skills.
  • Strong willingness to learn and develop technical skills.
  • Good communication and teamwork skills.
  • Self-motivated with a strong sense of ownership.

LigaData builds cutting-edge technology solutions tailored for data analytics, AI, and enterprise software, empowering clients throughout the Middle East and Africa. We’re focused on driving efficiency and insight for businesses looking to harness the full power of their data.

View company profile