AI academy - From Data Processing and Exploration to Learning Systems Without Supervision
This workshop connects the practical data-processing skills learned in the programme to the first family of machine learning methods that can learn without predefined labels.
The morning begins with a case-study keynote that frames data preparation as part of the learning process rather than a separate preliminary step. Participants revisit missing values, inconsistent categories, outliers, scaling, encoding, exploratory visualisation, and feature selection. They then examine why these choices directly shape distances, similarities, and the structure discovered by an unsupervised model.
The case study follows a mixed behavioural and service-use dataset in which the goal is to identify naturally occurring profiles without a known target variable. Participants explore distributions and relationships, prepare an analysis-ready feature space, and compare clustering approaches such as K-means and hierarchical clustering. Principal Component Analysis (PCA) is used to visualise high-dimensional structure and support interpretation.
Content
Morning
Keynote Lecture: From Raw Data to Structure Discovery
- Data quality, missingness, outliers, and analysis-ready data
- Exploratory Data Analysis (EDA) as a guide to model choice
- Feature scaling, encoding, and the geometry of similarity
- What unsupervised learning can and cannot infer
- K-means and hierarchical clustering
- Principal Component Analysis (PCA) for exploration and visualisation
- Interpreting clusters and avoiding overclaiming
Hands-on Lab: Discovering Structure in a Real-World Dataset
- Clean and prepare a mixed-feature dataset with Pandas
- Explore distributions and relationships visually
- Scale and select features for clustering
- Compare alternative cluster solutions
- Visualise
cluster structure with PCA
Afternoon
Group Case Study: From Clusters to Meaningful Profiles
- the analytical question and unit of analysis
- the selected features and preprocessing choices
- the preferred cluster solution and supporting evidence
- a plain-language interpretation of the discovered groups
- limitations
and what should be tested next
Group Presentations and Critical Review
Learning Outcomes
By the end of the course, participants will be able to:
- prepare and audit data for exploratory and unsupervised analysis
- explain how preprocessing choices influence distance-based learning methods
- apply clustering methods to identify structure in unlabeled data
- use PCA to explore and visualise high-dimensional patterns
- compare cluster solutions using quantitative and qualitative evidence
- interpret discovered groups critically and communicate their limitations
- formulate appropriate next-step questions from exploratory findings
Training Method
One-day intensive workshop combining:
- case-study keynote lecture
- guided hands-on coding in Jupyter notebooks
- data exploration and clustering exercises
- small-group analytical challenge
- presentations, peer review, and discussion
Certification
Certificate of ParticipationPrerequisites
Basic Python programming and familiarity with Pandas are required.
Planning and location
09:00 - 17:00