Skip to content
Skip to Content
Data & AI

AI academy - From Data Processing and Exploration to Learning Systems Without Supervision

This workshop connects the practical data-processing skills learned in the programme to the first family of machine learning methods that can learn without predefined labels.

The morning begins with a case-study keynote that frames data preparation as part of the learning process rather than a separate preliminary step. Participants revisit missing values, inconsistent categories, outliers, scaling, encoding, exploratory visualisation, and feature selection. They then examine why these choices directly shape distances, similarities, and the structure discovered by an unsupervised model.

The case study follows a mixed behavioural and service-use dataset in which the goal is to identify naturally occurring profiles without a known target variable. Participants explore distributions and relationships, prepare an analysis-ready feature space, and compare clustering approaches such as K-means and hierarchical clustering. Principal Component Analysis (PCA) is used to visualise high-dimensional structure and support interpretation.

In the hands-on lab, participants work through a guided Jupyter notebook to clean the data, standardise features, fit clustering models, compare alternative cluster solutions, and use internal measures such as inertia and silhouette score as supporting evidence rather than as automatic answers.
In the afternoon, groups receive a short analytical brief and must decide which features to include, how many clusters are defensible, what each cluster appears to represent, and what additional evidence would be needed before acting on the result. The day concludes with short presentations and a discussion on the limits of pattern discovery, including sensitivity to preprocessing, outliers, sampling, and interpretation.
Content
This workshop is structured into 3 progressive modules :

Morning
Keynote Lecture: From Raw Data to Structure Discovery
Topics covered through the keynote
  • Data quality, missingness, outliers, and analysis-ready data
  • Exploratory Data Analysis (EDA) as a guide to model choice
  • Feature scaling, encoding, and the geometry of similarity
  • What unsupervised learning can and cannot infer
  • K-means and hierarchical clustering
  • Principal Component Analysis (PCA) for exploration and visualisation
  • Interpreting clusters and avoiding overclaiming
Hands-on Lab: Discovering Structure in a Real-World Dataset
  • Clean and prepare a mixed-feature dataset with Pandas
  • Explore distributions and relationships visually
  • Scale and select features for clustering
  • Compare alternative cluster solutions
  • Visualise cluster structure with PCA

Afternoon
Group Case Study: From Clusters to Meaningful Profiles
Participants work in groups to develop and defend an unsupervised analysis pipeline. They define:
  • the analytical question and unit of analysis
  • the selected features and preprocessing choices
  • the preferred cluster solution and supporting evidence
  • a plain-language interpretation of the discovered groups
  • limitations and what should be tested next

Group Presentations and Critical Review
Short presentations (approximately 5 minutes per group)
Peer discussion on alternative interpretations and model sensitivity

Learning Outcomes

By the end of the course, participants will be able to:

  • prepare and audit data for exploratory and unsupervised analysis
  • explain how preprocessing choices influence distance-based learning methods
  • apply clustering methods to identify structure in unlabeled data
  • use PCA to explore and visualise high-dimensional patterns
  • compare cluster solutions using quantitative and qualitative evidence
  • interpret discovered groups critically and communicate their limitations
  • formulate appropriate next-step questions from exploratory findings

Training Method

One-day intensive workshop combining:

  • case-study keynote lecture
  • guided hands-on coding in Jupyter notebooks
  • data exploration and clustering exercises
  • small-group analytical challenge
  • presentations, peer review, and discussion
The training emphasises the link between data preparation, exploratory reasoning, unsupervised modelling, and responsible interpretation of discovered patterns.

Certification
Certificate of Participation
Prerequisites

Basic Python programming and familiarity with Pandas are required.

Participants should be comfortable with basic data visualisation and descriptive statistics. Prior experience with unsupervised learning is helpful but not required.

Planning and location
Session 1
08/09/2026 - Tuesday
09:00 - 17:00
Available Edition(s):
0.00 € 0.00 €

Your trainer(s) for this course