Zurück zu den Profis

Data Engineer

Überblick


Kurzer Überblick

Ein erfahrener Senior/Lead Data Engineer mit umfangreicher Erfahrung in verschiedenen Geschäftsbereichen, darunter Gesundheitswesen, Bildung, Web3.0, Finanzen, Medien, Werbung, Automobil, Telekommunikation, Luft- und Raumfahrt und Verteidigung. Beherrscht verschiedene Technologien wie Apache Spark, Apache Airflow, AWS, S3, Glue, Redshift, Lambda, Databricks, Snowflake, Scala, Python, Data Vault 2.0, Terraform und mehr. Zu seinen Leistungen gehören die Leitung von Teams, die Vertretung von Unternehmen bei der Interaktion mit Kunden, die Vorbereitung des technischen Designs, die Implementierung, die Codeüberprüfung und die erfolgreiche Durchführung von Projekten.

Berufliche Erfahrung

Co-founder and Lead Data Engineer

Business domain: Healthcare, Education, Web3.0, Finance

Start date: 2021-04-15

End date: Ongoing

Technologies: Apache Spark 3.0, Apache Airflow, AWS, S3, Glue, Redshift, Lambda, IAM, API Gateway, Databricks, Snowflake, DBT, Scala, Python, Data Vault 2.0, Terraform

Responsibilities:

  • Managing team of 7 data engineers, technical leadership and mentoring
  • Representing company in front of clients, analyzing business needs and proposing improvementsand technical guidance in the area of data engineering
  • Technical design preparation, implementation, code review, taking care of all aspects to make sure the solution can be delivered on time
  • Migration from on-premise Hadoop Cluster into the AWS and Snowflake based data lake for client from Healthcare industry.
  • Integration with Affiliate Platform used by Client. Preparing data flows and structures on top of AWS and Snowflake to identify customers which are coming from Affiliate Partners, track their activity to calculate fees for Affiliate Partners. I was leading the team of 5 developers and 1 QA tester. I was responsible for technical low-level design, implementation, code review and delivering the solution.
  • Migrating from Databricks Spark Streaming to near real time data flows built on Snowflake (using Streams, Hybrid Tables, Kafka Connector, Tasks). I was responsible for technical low-level design, implementation and verification during PoC phase and coordinating of all the work during production implementation phase.
  • Data Lake setup and Data ingestion for client from education domain (Web3). I was responsible for the setup, implementation, data lake design based on requirements from business users. We integrated company data from relational databases with blockchain (Solana) transaction details from Magic Eden and Solscan related to NFTs minted by the Client.

Senior Data Engineer

Business domain: Media and Advertisement

Start date: 2020-08-01

End date: 2021-04-30

Technologies: Apache Spark 3.0, Apache Airflow, Docker, Jenkins, AWS, Elastic MapReduce, S3, Glue, Redshift, Databricks, Data Lake, Scala, Python

Responsibilities:

  • Building ETL Pipelines in Apache Spark and Databricks / AWS EMR, Redshift, Glue, orchestrated on top of Airflow.
  • Performance tuning on Spark 3.0
  • Re-designing existing ETLs to cover all code best practises (unit test coverage, scalability, codequality)
  • ETL Migration from AWS EMR into Databricks environment
  • Design and implementation of CI/CD guidelines on top of Git and Jenkins

Senior Data Engineer

Business domain: Media and Advertisement

Start date: 2019-01-02

End date: 2020-08-01

Technologies: MapR, Apache Spark 2.3, Apache Kafka, Apache Airflow, HBase, Grafana, ELK, Docker, Openshift, Jenkins, Ansible, Python, Scala, SQL

Responsibilities:

- Building the platform based on MapR ecosystem that will be used to process data from test vehicles. The previous solution has been migrated from Cloudera cluster to the MapR environment.
  • Implementation of data flows using Apache Spark and performance tuning (the average daily volume of data for processing was close to 2PB of raw data every single day)
  • Preparing data flows on top of Hadoop Ecosystem orchestrated by Apache Airflow
  • Performance tuning of Apache Airflow deployment
  • Implementation of new features in Apache Airflow required internally (more advanced Spark Operators and Hooks, API authentication mechanism and many more)
  • Moving some data science applications to Openshift environment
  • Setting up Apache Airflow on top of Openshift / Kubernetes cluster using Celery Executor.Installation, environment preparation, security design, performance tuning and scaling.
  • GPS Data processing and visualization using Apache Spark and ELK stack (Elasticsearch, Kibana)

Software Engineer / Senior Software Engineer

Business domain: Automotive, Telecommunications, Aerospace and Defence

Start date: 2015-11-02

End date: 2018-12-31

Technologies: Apache Spark, Apache Hive, Apache Impala, Apache Hadoop, Apache Kafka, Cloudera, Hortonworks, Python, Scala, SQL, AWS

Responsibilities:

  • Preparing data flows on top of Hadoop Environment (Spark, Sqoop, Oozie, Hive, Pig, Shell)
  • Creating and performance tuning of Apache Spark applications
  • Near Real-Time ETL solution implementation (Apache Kafka, Apache Spark- Structured Streaming, Apache Hive LLAP)
  • Analysing signals from different modules in cars to find answers for science questions via HiveQL- Queries.
  • Conducting multiple trainings for internal employees (Big Data Architecture, Apache Hive, Apache Spark)
  • Sharing knowledge internally in communities (Big Data & Data Science)
  • Implementation of multiple applications which help to create Big Data Solutions faster and better(Oozie workflow generator, Metadata management)
  • Cleaning and preprocessing data (Python / Spark)
  • Building and applying clustering algorithms (Scikit learn)
  • Implementing Autoencoder Neural Network model in Tensorflow and Keras on GPU virtual machine
  • Mentorship and technical leadership for junior developers

CRM Analyst

Technologies:SQL, Python, Teradata, SAS, Hadoop, Control-M

Start date: 2014-09-01

End date: 2015-10-30

Responsibilities

- Developing new analytical and operational CRM platform
  • Implementing inbound and outbound, multichannel and real-time marketing campaigns in SASCustomer Intelligence Studio
  • Sentiment analysis using data science and text mining Python libraries
  • Testing algorithm which is used to categorize transactions based on MCC codes

English Level – C1 German Level – B1

Author
Data Engineer
653a0b0bb55a4a0001411098
Verfügbarkeit06/11/2023
Freiberufler440  / Pro Tag
EbenSENIOR
Rollen
Data Engineer
Sprachen
English
Fähigkeiten
AI and Machine LearningAWSAnsibleApache KafkaApache SparkAutomation SoftwareBig DataBlockchainCRMCertificationsCloudClouderaClustersCode ManagementContainer Engine SoftwareContainer Management SoftwareContainerization SoftwareDatabaseDatabricksDevOpsDockerElasticsearchGitGrafanaHadoopKerasKibanaKubernetesMessage BrokerMonitoring ToolsOpenShiftProject ManagerPythonSQLSW Tools APIsScalaScikit LearnSearch ToolsShellSnowflakeStream Analytics SoftwareTensorFlowTeradataTerraformTesterTools
Lokalitäten
Remote