This course sits at the intersection of Data Science and Artificial Intelligence.
It is designed for future business leaders who want to engage with data directly rather than defer to the analysts in the room.
Its primary objectives are to:
- become fluent in the language of “data science and AI”,
- learn how to effectively communicate uncertainty,
- cultivate the ability to craft compelling narratives grounded in data,
-
use data analytics for informed decision-making in uncertain environments,
-
develop skills to explore and make sense of large and messy data,
-
build and interpret (predictive) models with confidence,
-
learn to direct AI coding with Claude Code to produce a defensible data analysis.
Ideal for students preparing for careers in data-rich environments, the course emphasizes data storytelling. The goal is to learn how to turn data into a clear, honest narrative a decision-maker can act on.
Each lecture features two to three authentic datasets, whose analysis is demonstrated live in class. Examples range across consumer database mining, internet and social media tracking, asset pricing, network analysis, healthcare, sports analytics, and text mining, so the methods always arrive attached to a real business decision. For example, we will analyze a social network of the Medici family in Florence, decode who wrote the Federalist, build song recommendation systems, understand the demographics of Titanic survivors, or determine biomarkers for leukemia.
The curriculum spans topics from classical statistics (e.g. hypothesis-driven decisions), data science (dimensionality reduction) to modern machine learning techniques (e.g deep learning). It also explores cutting-edge advancements in generative AI. The course puts a particular emphasis on the analysis of text data in the context of both small and Large Language Models (LLM) that form a basis of popular text-generating systems. Techniques covered include large-scale testing and false discovery rates, modern regression and model choice, machine-learning based classification, network analysis, language and topic models, principal components, clustering, Bayesian analysis, deep learning, transformers and attention.
By the end of the course, students will be equipped to perform machine-supported intelligent data analysis and communicate findings effectively.