Data Pipeline, Data Ingestion, Data Manipulation, Cleaning Data and Schedule Data - Engineering Workflows (classroom, synchronous, asynchronous e-learning)
About This Course
Note: This course is offered as a module in the (SCTP) Junior Data Engineer programme. We do not offer this module on its own. Please scroll down to proceed to the main SCTP site for more information and to register interest.
This course curriculum prepares entry level professionals to handle the data management activities such as collecting the data from internal and external data sources (big data), building the data structure to store the data captured, and loading the data into the data systems (such as data lake) as well as have a good understanding in different types of database systems such as traditional relational databases and modern big data management platforms, and the data transformation techniques. There are behavioural skills and mindsets elements integrated into both the technical and practice sessions.
What You'll Learn
- the foundations of a data platform and explore different forms of testing for a data transformation pipeline
- how to import and ingest data from various sources
- the way to manipulate DataFrames as well as time series data using pandas in Python
- the terminology, methods and processes to prepare data processes using Python with Apache Spark
- how to clean data stored in an SQL Server database
- how to schedule data engineering workflows using Airflow in Python. They will then apply these skills to build a production-quality workflow
Throughout the module, participants will get hands-on practice with all tasks using a wide range of interesting and messy datasets on DataCamp.
Entry Requirements
Pass screening which includes a basic technical test and interview