Submitting more applications increases your chances of landing a job.

Here’s how busy the average job seeker was last month:

Opportunities viewed

Applications submitted

Keep exploring and applying to maximize your chances!

Looking for employers with a proven track record of hiring women?

Click here to explore opportunities now!
We Value Your Feedback

You are invited to participate in a survey designed to help researchers understand how best to match workers to the types of jobs they are searching for

Would You Be Likely to Participate?

If selected, we will contact you via email with further instructions and details about your participation.

You will receive a $7 payout for answering the survey.


User unblocked successfully
Thank you. Your report has been submitted and will be reviewed shortly.
Ahmed Mab, Senior Data Engineer – Java, PySpark

Ahmed Mab

Senior Data Engineer – Java, PySpark·MGEN

Tunisia

Master's degree, Statistical Engineering – Statistics and Computer Science

Work experience

Total years of experience: 11 years, 6 months

Senior Data Engineer – Java, PySpark

July 2021 - July 2025

MGEN

Paris, France

July 2021 - July 2025

Worked as a Data Engineer on the design and development of data processing applications for MGEN, a major French health insurance organization.

The platform was designed to ingest, transform, enrich, store, and expose data related to members, insurance contracts, and healthcare benefits.

Main responsibilities:

• Designed and developed data processing applications using Java, PySpark, Couchbase, and cloud technologies.

• Built ingestion pipelines to extract and load data from text, CSV, and JSON files into object storage.

• Developed Java-based transformation processes to cleanse, aggregate, normalize, and enrich large volumes of business data.

• Loaded and persisted processed data in Couchbase and cloud-based storage environments.

• Extracted, processed, and transformed data from Greenplum databases.

• Designed and developed RESTful web services using Java, Spring Boot, and Spring Data to expose data stored in Couchbase to downstream applications.

• Developed PySpark jobs for log processing, analysis, and monitoring.

• Indexed application and processing logs in Elasticsearch and contributed to the creation of operational monitoring dashboards.

• Worked with Databricks for distributed data processing and data engineering activities.

• Used Linux Shell scripting to automate data processing and operational tasks.

• Participated in Agile ceremonies, technical design, code reviews, testing, troubleshooting, and production support.

• Managed source code and collaborative development workflows using Git and GitLab.

Technical environment: Java, Python, PySpark, Apache Spark, Couchbase, Spring Boot, Spring Data, Elasticsearch, Greenplum, Databricks, Google Cloud Platform, S3-compatible object storage, Linux, Shell scripting, Git, GitLab, Jira, Agile.

Company industry:
Insurance & TPA

Senior Big Data Engineer – Spark / PySpark

January 2020 - July 2021

ACOSS

Montreuil, France

January 2020 - July 2021

Worked as a Senior Big Data Engineer on the migration and modernization of the French Nominal Social Declaration reporting system, known as DSN, within a distributed Big Data environment.

Main responsibilities:

• Contributed to the migration of the DSN business intelligence and reporting platform to a scalable Big Data architecture.

• Designed and developed large-scale data ingestion and transformation applications using Python, PySpark, Java 8, and Apache Spark.

• Developed Spark jobs to ingest, process, transform, and enrich large volumes of social declaration data.

• Designed optimized data storage models in Apache HBase and made processed datasets available to business users and downstream applications.

• Improved data storage, access, and processing performance within the distributed Hadoop ecosystem.

• Used the Spark-Kafka connector to consume data from Kafka topics and publish processed data to downstream topics.

• Integrated Spark processing jobs with Apache Kafka, HDFS, Hive, and HBase.

• Worked in a secured Kerberos-enabled Hadoop environment.

• Participated in Agile ceremonies, code reviews, testing, version control, and continuous delivery activities.

Technical environment: Python, PySpark, Java 8, Apache Spark, Apache Kafka, Hadoop, HDFS, Apache Hive, Apache HBase, Kerberos, Git, GitLab, Jira, Agile.

Company industry:
Public Administration

Senior Big Data Engineer – AWS & Spark

December 2018 - December 2019

Believe Distribution Services

Paris, France

December 2018 - December 2019

Worked as a Senior Cloud Data Engineer on the development of distributed data ingestion and transformation pipelines for Believe Distribution Services.

Designed and developed large-scale data processing pipelines using Apache Spark in a distributed AWS cloud environment.

Main responsibilities:

• Developed data ingestion and transformation pipelines using Spark, PySpark, and Scala.

• Processed, standardized, and normalized daily music consumption and sales reports received from digital streaming platforms, including Spotify, Apple Music, Deezer, Napster, and iTunes.

• Implemented data validation, cleansing, transformation, and enrichment processes to improve data quality and consistency.

• Enriched streaming and batch datasets using reference data stored in Amazon S3.

• Executed distributed Spark workloads using Amazon EMR and AWS processing steps.

• Developed Spark DataFrame-based processing jobs for large-scale music and artist data.

• Implemented machine learning algorithms using Spark ML to classify and categorize artists.

• Developed real-time and near-real-time data processing pipelines using Spark Structured Streaming.

• Implemented serverless processing components using AWS Lambda.

• Managed application and processing logs using the ELK stack.

• Integrated data pipelines with MariaDB, HDFS, Amazon S3, Apache NiFi, and downstream information systems.

• Containerized data applications using Docker and contributed to deployment and version-control processes using Git and GitLab.

• Worked with Databricks and Azure Machine Learning for data processing and machine learning experimentation.

Technical environment: Python, PySpark, Scala, Apache Spark, Spark DataFrames, Spark ML, Spark Structured Streaming, AWS, Amazon EMR, Amazon S3, AWS Lambda, EMR Steps, Hadoop, HDFS, Apache NiFi, MariaDB, ELK Stack, Docker, Databricks, Azure Machine Learning, Git, GitLab.

Company industry:
Media Production

Senior Data Engineer – Big Data / PySpark

April 2018 - December 2018

AXA France

Nanterre, France

April 2018 - December 2018

Worked as a Senior Data Engineer on Big Data projects within AXA France’s industrialized data platform.

Main responsibilities:

• Designed and developed scalable data processing applications using Python and PySpark.

• Defined and implemented reusable data integration patterns for ingesting data into the enterprise data lake.

• Managed the complete data lifecycle, from raw data ingestion to transformation, enrichment, validation, and business data exposure.

• Implemented data quality controls to ensure the accuracy, consistency, completeness, and reliability of processed datasets.

• Built data pipelines to transform raw data into curated business views for reporting, analytics, and operational use cases.

• Integrated processed data and analytical results back into AXA’s legacy information systems.

• Automated and orchestrated batch data processing workflows using Apache Oozie, Jenkins, and Shell scripting.

• Worked in an industrialized development environment using Git, GitLab, Jira, and continuous integration practices.

Technical environment: Python, PySpark, Apache Spark, Hadoop, HDFS, Apache Oozie, Jenkins, Shell scripting, Git, GitLab, Jira.

Company industry:
Insurance & TPA

Big Data Tech Lead – Spark & Kafka

April 2017 - April 2018

EDF

La Defense, France

April 2017 - April 2018

Worked as a Big Data Tech Lead on a real-time streaming and monitoring solution for EDF’s Customer Commerce Division.

Designed the architecture and contributed to the development of a Big Data streaming platform used to monitor the activities of customer service advisors across EDF contact centers. The solution helped improve customer service response times and quality, contributing to higher customer satisfaction and lower customer attrition.

Key scale and performance indicators:
• Collected and processed more than 500, 000 events per day generated by approximately 150, 000 incoming calls.
• Monitored Front Office and Back Office activities for more than 10, 000 customer service advisors.
• Ensured end-to-end traceability of customer interactions across phone, mail, email, web calls, and other communication channels.
• Delivered real-time data with a response latency of less than one second.

Main responsibilities:
• Designed and developed a reusable Core module shared across all project jobs, covering configuration management, standardized logging, and access to reference-data connectors.
• Developed a Kafka consumer using Spark Streaming to enrich, transform, and calculate streaming data before continuously indexing it into Elasticsearch.
• Configured and developed Apache Oozie coordinators and workflows.
• Developed a daily batch process to replicate reference tables from Oracle databases to HDFS.
• Designed and implemented Kibana dashboards and real-time operational visualizations.
• Created unit tests and mocks to simulate Elasticsearch, Kafka, and Spark contexts.
• Supported the production operations team with job deployment and automation using Jenkins and Ansible.
• Provided technical guidance and contributed to development standards and solution architecture.

Technical environment: Java 8, Apache Kafka 0.10, Spark Streaming, Elasticsearch 6 with X-Pack, Hortonworks, Hadoop, HDFS, Kerberos, Apache Oozie, Apache NiFi, Oracle, Kibana, Jenkins, Ansible, Git, GitLab, JUnit, Maven, Nexus, Jira, Confluence, Shell scripting, Node.js.

Company industry:
Electric Power Production & Transmission

Big Data Engineer / Spark Developer

February 2016 - March 2017

ENEDIS

Nanterre, France

February 2016 - March 2017

Worked as a Big Data Developer on data processing and analytics projects for ENEDIS, the French electricity distribution network operator.

• Designed and developed an Apache Spark application using Java to support France’s electricity capacity mechanism, helping ensure electricity supply security during winter peak-demand periods.

• Developed RESTful web services using Java and Dropwizard to retrieve, process, and expose electricity load curve data stored in Elasticsearch.

• Designed and implemented a Big Data platform for analyzing electric vehicle charging loads and usage data.

• Developed and integrated distributed data processing pipelines using Spark, PySpark, Hadoop, Kafka, NiFi, Teradata, and Shell scripting.

Technical environment: Apache Spark, Java, PySpark, Hadoop, Elasticsearch, Teradata, Apache Kafka, Apache NiFi, Dropwizard, Shell scripting.

Company industry:
Electric Power Production & Transmission

Big Data & Machine Learning Developer

February 2014 - March 2016

Lansrod Consulting

Tunis, Tunisia

February 2014 - March 2016

Worked as a Big Data Developer on data collection, real-time analytics, and machine learning projects.

Main responsibilities:

• Developed an online reputation monitoring platform to collect, process, transform, and index large volumes of data from web sources.

• Built data ingestion and processing pipelines using Apache NiFi, Kafka, Spark, Hadoop, and MapReduce.

• Implemented a real-time log collection and analytics platform using Kafka and the ELK Stack.

• Developed data processing applications using Java, Python, Spark, Hive, and Hadoop.

• Implemented machine learning algorithms to analyze online user opinions and perform sentiment analysis.

• Created Hadoop workflows and scheduled data processing jobs using Apache Oozie.

• Used Hive and Hue to query, explore, and analyze large datasets stored in the Hadoop ecosystem.

• Contributed to application development and data visualization components using Node.js.

• Worked in an Agile development environment using Git and Jira.

Technical environment: Apache Spark, Apache NiFi, Kafka, ELK Stack, Elasticsearch, Logstash, Kibana, Hadoop, MapReduce, Hive, Oozie, Hue, Java, Python, Node.js, Git, Jira.

Company industry:
IT Services

Education

École Supérieure de la Statistique et de l’Analyse de l’Information (ESSAI)

June 2014

June 2014

Master's degree, Statistical Engineering – Statistics and Computer Science

Tunisia

Institut Préparatoire aux Études d’Ingénieurs de Monastir (IPEIM)

June 2010

June 2010

Higher diploma, Preparatory Studies for Engineering Schools – Mathematics and Technology

Tunisia

Lycée Eljem

June 2008

June 2008

Bachelor's degree, Mathematics and Technical Sciences

Tunisia

Languages

Arabic

Native Speaker

French

Expert

English

Intermediate