Measure Theory, Probability, and Stochastic Processes
Page Contents
ToggleCourse Overview
This page brings together graduate-level lectures and exercises on measure theory, probability, and stochastic processes. The material follows a unified measure-theoretic approach, structured in three parts: Part I — Measure Theory and Integration, Part II — Probability and Random Sequences, and Part III — Stochastic Processes and Ergodic Theory.
Each lecture is available as a downloadable PDF, accompanied by exercises with detailed solutions. Selected lectures also include video recordings in English and French.
This foundational framework underpins the resources available on Statistics Theory, Machine & Deep Learning, Information Theory, Queueing Theory, and Point Processes.
Part I: Measure Theory and Integration
1 Measure Theory
PDF, video, and exercises of the lecture on Measure Theory, the cornerstone of modern integration and probability theory.
We begin with the fundamental concept of measurability, exploring algebras and σ-algebras, the construction of generated and Borel σ-algebras, product σ-algebras, and the Borel structure of Euclidean and extended spaces. We then study measurable functions, with criteria for measurability, its stability under operations and limits, the notion of simple functions and the indispensable simple approximation theorem, together with semicontinuity and the functional monotone class theorem. We move on to a rigorous examination of measures, from their definition and basic properties — illustrated by the Dirac measure and the weighted counting measure — to the notions of finite, probability, and σ-finite measures, the continuity of measures, and the extension and uniqueness of measures from algebras and semirings. We conclude with distribution functions and the Lebesgue measure: cumulative distribution functions, the concept of almost everywhere, and the existence and uniqueness of the Lebesgue measure on ℝⁿ, characterized by its translation invariance (Haar measure).
Course Outline:
- 1.1 Notations
- 1.2 Measurability of Sets
- Measurable Sets and σ-Algebra
- Product σ-Algebra
- Borel σ-Algebra on Euclidean and Extended Spaces
- 1.3 Measurability of Functions
- Measurable Functions
- Simple Approximation Theorem
- Upper and Lower Semicontinuity
- Functional Monotone Class Theorem
- 1.4 Definition and Extension of Measures
- Definition and Basic Properties
- Extension and Uniqueness
- Semirings and Measure Extension
- 1.5 Distribution Functions and the Lebesgue Measure
- Cumulative Distribution Function (CDF)
- μ-Almost Everywhere
- Lebesgue Measure
- Translation Invariance and Haar Measure
2 Lebesgue Integral
PDF, video, and exercises of the lecture on the Lebesgue Integral, a profound extension of the classical Riemann integral.
We begin with the construction of the Lebesgue integral in three stages — simple, nonnegative, then integrable functions — and establish its fundamental properties and comparison criteria. The heart of the theory is the exchange of integral and limit, through the Monotone Convergence Theorem, Fatou’s Lemma, and the Dominated Convergence Theorem, with applications to series as Lebesgue integrals, the change of variable formula, and integrals depending on a parameter. We then turn to Fubini’s theorem and the Fubini-Tonelli theorem, governing the interchange of the order of integration, together with integration by parts for measures and CDFs. We present the Radon-Nikodym theorem, central to the relationship between measures, with the notions of absolute continuity and the Lebesgue decomposition of measures. We then develop the ℒp and Lp spaces: Minkowski’s and Hölder’s inequalities, the Lp norm and its completeness (the Riesz-Fischer theorem), essential boundedness, and the Hilbert space L2. We extend the theory to vector-valued Lp spaces, covering the vectorial integral, locally integrable functions, the smooth change of variables, and the Lebesgue Differentiation Theorem. We introduce the support of a measure and the essential support of a function, and conclude with the dense subsets in Lp.
Course Outline:
- 2.1 Construction and Properties of the Integral
- Construction of the Lebesgue Integral
- Properties of the Lebesgue Integral
- Comparison Criteria by Inequality
- Monotone and Dominated Convergence Theorems
- Series as Lebesgue Integral
- Change of Variable in Integrals
- Integrals Depending on a Parameter
- 2.2 Fubini and Radon-Nikodym Theorems
- Fubini’s Theorem
- Integration by Parts for Measures and CDFs
- Radon-Nikodym Theorem
- Lebesgue Decomposition of Measures
- 2.3 ℒp and Lp Spaces
- ℒp Spaces
- Minkowski and Hölder Inequalities
- Lp Spaces
- Lp Norm and Completeness
- Essential Boundedness and First Mean Value
- 2.4 Vector-Valued Lp Spaces
- Vectorial Integral
- Vectorial Lp Space
- Locally Integrable Functions
- Smooth Change of Variables
- Lebesgue Differentiation Theorem
- 2.5 Support of a Measure and a Function
- Support of a Measure
- Essential Support of a Function
- 2.6 Dense Subsets in Lp
3 Integration of Functions of a Real Variable
PDF, video, and exercises of the lecture on Integration of Functions of a Real Variable, adapting to the Lebesgue framework classical results that are traditionally presented for the Riemann integral.
After constructing the Lebesgue integral on general measure spaces in the previous lectures, this chapter restricts to mappings of a real variable, with ℝ equipped with the Lebesgue measure. Many classical results of integration on ℝ are traditionally stated for the Riemann integral, as in the well-known textbook Cours de Mathématiques Spéciales by Ramis, Deschamps and Odoux. We adapt them here to the Lebesgue framework, for mappings with values in ℝ̄n or ℂn. We develop the Chasles relation, integral mapping and antiderivatives, the mean value formulas, the Newton-Leibniz formula, integration by parts, change of variables, Taylor’s formula with integral remainder, integrability by comparison, the integration of comparison relations, and finally the improper integral with its calculation techniques. A pedagogical highlight is that integration is a regularizing operation: starting from a possibly discontinuous integrand, the integral mapping is always continuous, and differentiable almost everywhere thanks to the Lebesgue Differentiation Theorem.
Course Outline:
- 3.1 Properties of the Integral on ℝ
- Chasles Relation
- Integral Function
- Antiderivatives
- Mean Value Formulas
- 3.2 Integral vs. Antiderivatives
- Newton-Leibniz Formula
- Integration by Parts
- Change of Variables
- Taylor’s Formula with Integral Remainder
- 3.3 Integrability by Comparison
- Integrability by Inequality
- Integrability by Domination and Equivalence
- 3.4 Integration of Comparison Relations
- Comparison Relations: Integrable Case
- Comparison Relations: Non-Integrable Case
- Integrability Rules
- 3.5 Improper Integral
- Improper Integral by Passing to the Limit
- Properties of Improper Integrals
- Abel’s Rule and Fundamental Examples
- Bilateral Improper Integral
- 3.6 Calculation Techniques for an Improper Integral
- Use of Antiderivatives
- Change of Variables for the Improper Integral
- Integration by Parts
Part II: Probability and Random Sequences
4 Probability Theory
PDF, video, and exercises of the lecture on Probability Theory, developed from its measure-theoretic foundations.
We begin with the probability space, a measure space of total mass one: the probability measure, events, the notion of « almost surely », and conditional probability. We then study random variables and their distributions, with the expectation as a Lebesgue integral — via the transfer theorem — followed by moments, the variance, the characteristic function, and the cumulative distribution and probability density functions (CDF and PDF). We devote a section to independence: independent sets, families, and random variables, the product law and product expectation, the additivity of the variance under independence, and the change of variables in multivariate distributions. Finally, we examine the monotone and dominated convergence theorems for random variables, transposing the cornerstone limit theorems of integration into the probabilistic setting.
Course Outline:
- 4.1 Probability Space
- 4.2 Random Variable
- Expectation as Integral
- Moments and Characteristic Function
- Cumulative Distribution Function (CDF) and Probability Density Function (PDF)
- 4.3 Independence, Product Law, and Change of Variables
- Independent Sets, Families, and Random Variables
- Product Law, Product Expectation, and Additivity of the Variance
- Change of Variables in Multivariate Distributions
- 4.4 Monotone and Dominated Convergence for Random Variables
- Monotone Convergence for Random Variables
- Dominated Convergence for Random Variables
5 Discrete Random Variables and their Transform
PDF and exercises of the lecture on Discrete Random Variables and Their Transform.
We begin with the four fundamental discrete distributions — the Bernoulli, binomial, geometric, and Poisson random variables — each with its distribution on the support, its mean, and its variance. We then introduce the probability generating function (PGF), the power series that encodes the distribution of an ℕ-valued random variable: its radius of convergence, the recovery of the distribution from its derivatives at the origin, the generating function of a sum of independent variables, and the computation of factorial moments. Finally, we study random sums of random variables — obtained by composing generating functions, with Wald’s formula for the expectation — and the monotonicity and convexity of the generating function, characterizing the fixed points of g(x) = x.
Course Outline:
- 5.1 Fundamental Discrete Random Variables
- Bernoulli Random Variable
- Binomial Random Variable
- Geometric Random Variable
- Poisson Random Variable
- 5.2 Probability Generating Function
- Examples of Probability Generating Functions
- Probability Generating Function and Distribution Analysis
- Factorial Moments from Probability Generating Function
- 5.3 Random Sums and Convexity of the Generating Function
- Random Sum of Random Variables
- Monotonicity and Convexity of the Probability Generating Function
6 Continuous Random Variables and their Transforms
PDF and exercises of the lecture on Continuous Random Variables and Their Transforms.
We study the fundamental continuous random variables — the Gaussian random variable, defined by its bell-shaped density and characterized by its mean and variance, and the exponential random variable, distinguished by its lack-of-memory property. We then develop the three integral transforms that encode a distribution through its moments. The moment generating function (MGF) is introduced through its definition and its radius of convergence, followed by its infinite series expansion, from which each moment is recovered as a derivative at the origin. The characteristic function — the Fourier transform of the distribution — always exists and is bounded and uniformly continuous; we cover its differentiability and finite expansion, the pivotal Lévy’s inversion formula that recovers the distribution from the transform, and its infinite expansion, real analytic whenever the radius of convergence is positive. Finally, the Laplace transform, tailored to nonnegative random variables where it is always defined, is treated through its differentiability and finite expansion, its infinite expansion, and its real analyticity. Each of the three transforms recovers the moments of the distribution from its derivatives at the origin and characterizes it.
Course Outline:
- 6.1 Fundamental Continuous Random Variables
- Gaussian Random Variable
- Exponential Random Variable
- 6.2 Moment Generating Function
- Definition and Radius of Convergence
- Infinite Expansion of the Moment Generating Function
- 6.3 Characteristic Function
- Characteristic Function: First Properties
- Differentiability and Finite Expansion of Characteristic Function
- Lévy’s Inversion Formula
- Infinite Expansion of the Characteristic Function
- 6.4 Laplace Transform
- Laplace Transform: Definition
- Differentiability and Finite Expansion of Laplace Transform
- Infinite Expansion of the Laplace Transform
- Laplace Transform is Real Analytic
7 Random Vectors and Gaussian Distribution
PDF and exercises of the lecture on Random Vectors and Gaussian Distribution.
We begin with the characteristic function of a random vector, with Lévy’s inversion formula for random vectors — pivotal for recovering the distribution from the characteristic function — and the independence criterion that determines the statistical independence of components. We then study square-integrable random vectors, introducing the covariance matrix (alongside the pseudo-covariance matrix), the notion of degenerate random vectors whose covariance matrix is non-invertible, and the effect of affine transformations on mean and covariance. We turn to Gaussian random vectors, covering their definition and characteristic function, the equivalence between independence and uncorrelation for jointly Gaussian vectors, and the probability density function of a nondegenerate Gaussian vector. We next develop complex Gaussian random vectors: the basics of complex random vectors, symmetric (circularly symmetric) complex Gaussian vectors denoted 𝒞𝒩, the independence criterion for jointly 𝒞𝒩 vectors, and their spectral decomposition and characterization. Finally, we cover the real representations of complex vectors and matrices, the covariance and linear combinations of 𝒞𝒩 random vectors, and the probability density function of 𝒞𝒩 random vectors.
Course Outline:
- 7.1 Characteristic Functions of Random Vectors
- Lévy’s Inversion Formula for Random Vectors
- Independence Criterion
- 7.2 Square-Integrable Random Vectors
- Covariance Matrix
- Degenerate Random Vectors
- Affine Transformations
- 7.3 Gaussian Random Vectors
- Definition and Characteristic Function
- Independence Versus Uncorrelation for Gaussian Vectors
- Probability Density Function of a Nondegenerate Gaussian Vector
- 7.4 Complex Gaussian Random Vectors
- Basics of Complex Random Vectors
- Symmetric Complex Gaussian Vector
- Criterion for Independence of Jointly 𝒞𝒩 Random Vectors
- Spectral Decomposition and Characterization of 𝒞𝒩 Vectors
- 7.5 Real Representations and Density of Complex Gaussian Vectors
- Real Representations of Complex Vectors and Matrices
- Covariance and Linear Combinations of 𝒞𝒩 Random Vectors
- Probability Density Function of 𝒞𝒩 Random Vectors
8 Convergence for Sequences of Random Variables
PDF and exercises of the lecture on Convergence for Sequences of Random Variables.
We begin with almost-sure convergence, covering the Strong Law of Large Numbers, the Borel-Cantelli lemmas, and various conditions guaranteeing almost-sure convergence. We then turn to convergence in probability, a weaker mode, with the Continuous Mapping Theorem for convergence in probability. We treat convergence in quadratic mean (L² convergence), with the Cauchy criterion and the continuity of the inner product. We progress to weak convergence (convergence in distribution), covering the Portmanteau theorem, the Continuous Mapping Theorem for weak convergence, the Characteristic Function Criterion, and the fundamental Central Limit Theorem. We then synthesize the modes through the connections between convergence types, introduce stochastic order for comparing the magnitudes of random variables, and address the tightness of probability measures with the pivotal Prohorov’s Theorem.
Course Outline:
- 8.1 Almost-Sure Convergence
- Strong Law of Large Numbers
- Borel-Cantelli Lemmas
- Conditions for Almost-Sure Convergence
- 8.2 Convergence in Probability
- Continuous Mapping Theorem for Convergence in Probability
- 8.3 Convergence in Quadratic Mean
- 8.4 Weak Convergence — Convergence in Distribution
- Portmanteau Theorem
- Continuous Mapping Theorem for Weak Convergence
- Characteristic Function Criterion
- Central Limit Theorem
- 8.5 Connections Between Convergence Types
- 8.6 Stochastic Order
- 8.7 Tightness of Probability Measures
- Tightness and Uniform Tightness
- Lemmas on Tightness
- Prohorov’s Theorem
Part III: Stochastic Processes and Ergodic Theory
9 Stochastic Processes
PDF and exercises of the lecture on Stochastic Processes, mathematical models depicting the evolution of systems over time through probabilistic mechanisms.
We begin by defining stochastic processes and the notion of independence within them, then develop finite-dimensional distributions (fidis), Kolmogorov’s Extension Theorem — which constructs a process from its finite-dimensional distributions — and the analysis of sample paths. We turn to second-order processes, with their mean and covariance functions, and to stopping times for discrete-time processes. We then distinguish strict stationarity, which requires the joint distribution of any set of points to be invariant under time shifts, from wide-sense stationarity, which relaxes this to invariance of the mean and autocovariance alone. We cover the measurability of stochastic processes and the stochastic integral. Finally, we develop Gaussian and Wiener processes: the definition and characteristic function of a Gaussian process, the conditions for its existence and stationarity, Gaussian Hilbert subspaces — where the closed subspace generated by a Gaussian family remains Gaussian — and the Wiener process, also known as Brownian motion, a cornerstone of stochastic modeling.
Course Outline:
- 9.1 Foundations of Stochastic Processes
- Definition and Independence of Stochastic Processes
- Finite-Dimensional Distributions (fidis)
- Kolmogorov’s Extension Theorem
- Sample Paths of a Stochastic Process
- 9.2 Second-Order Processes and Stopping Times
- Second-Order Stochastic Process
- Stopping Times
- 9.3 Strict and Wide-Sense Stationarity
- Stationary Stochastic Processes
- Wide-Sense Stationarity
- 9.4 Measurability and Stochastic Integral
- Measurable Stochastic Processes
- Stochastic Integral
- 9.5 Gaussian and Wiener Processes
- Definition and Characteristic Function of Gaussian Process
- Existence and Stationarity of Gaussian Process
- Gaussian Hilbert Subspaces
- Wiener Process — Brownian Motion
10 Conditional Expectation and MMSE Estimation
PDF and exercises of the lecture on Conditional Expectation and MMSE Estimation, two fundamental concepts in probability theory and statistical estimation.
We begin with conditional expectation, first conditioning with respect to a σ-algebra and then with respect to a random variable. We develop its properties and the monotone and dominated convergence theorems for conditional expectation. We then treat the special cases: conditional expectation for discrete, continuous, and mixed variables, and the Gaussian case, covering both real jointly Gaussian vectors and complex 𝒞𝒩 vectors. We turn to Minimum Mean Square Error (MMSE) estimation, showing that the conditional expectation is precisely the MMSE estimator, and finally to linear MMSE estimation for random vectors with invertible observation covariance.
Course Outline:
- 10.1 Conditional Expectation
- Conditioning With Respect to a σ-Algebra
- Conditioning With Respect to a Random Variable
- Properties of Conditional Expectation
- Monotone and Dominated Convergence for Conditional Expectation
- 10.2 Conditional Expectation in Special Cases
- Conditional Expectation for Discrete and Continuous Variables
- Conditional Expectation in Gaussian Case
- 10.3 MMSE Estimation
- MMSE Estimation vs. Conditional Expectation
- Linear MMSE Estimation
11 Martingales and Ergodic Theory
PDF and exercises of the lecture on Martingales and Ergodic Theory, two pivotal areas in probability and dynamical systems.
We begin with an introduction to martingales, focusing on their definition and basic properties, highlighting their significance in stochastic processes. We delve into the optional stopping theorem and martingale convergence theorems, crucial for understanding the behavior of martingales under various conditions, and examine several examples of martingales illustrating their application in different probabilistic contexts. We then turn to ergodic theory, starting with the notions of flow, shift, and compatibility, essential for understanding the dynamics of systems over time. We introduce the stationary framework, providing a foundation for analyzing systems that exhibit statistical regularity. A key highlight is Birkhoff’s Pointwise Ergodic Theorem, a fundamental result that connects individual trajectories with long-term statistical behavior. Finally, we explore the ergodic stationary framework, offering insights into the behavior of systems that are both stationary and ergodic.
Course Outline:
- 11.1 Martingales
- Definition and Basic Properties of Martingales
- Optional Stopping and Martingale Convergence Theorems
- Examples of Martingales
- 11.2 Ergodic Theory
- Flow, Shift and Compatibility
- Stationary Framework
- Birkhoff’s Pointwise Ergodic Theorem
- Ergodic Stationary Framework
Book Coming Soon
A book based on this material is currently in preparation:
Mohamed Kadhem KARRAY — « Measure Theory, Probability, and Stochastic Processes: An Integrated Approach ».
The book will provide a unified treatment of measure theory, probability, and stochastic processes, expanding and consolidating the lectures available on this page. Stay tuned for the publication announcement.
Acknowledgements
This book grew out of my collaboration with François Baccelli and Bartłomiej Błaszczyszyn on our co-authored « Random Measures, Point Processes, and Stochastic Geometry » [2], which led me to revisit the foundations of measure theory and probability in greater depth. The ergodic theory in the Martingales and Ergodic Theory chapter builds on the ergodicity material we developed together there, within the stationary-process framework of Baccelli and Brémaud [3]. I am grateful to Bartłomiej for a long-standing collaboration that has borne much fruit in measure and probability theory.
I also thank my PhD student, Lucas Darlavoix, for our joint work during his thesis [21], which connects to the Measure Theory, Lebesgue Integral, and Probability Theory chapters, and with whom I co-authored a book [38] applying these foundations to data science and wireless networks.
I owe a deep debt to Pierre Brémaud, whose contributions to probability theory [10, 13, 15, 16, 17] have profoundly shaped this book. Beyond his published works, I am especially indebted to the unpublished lecture notes he developed for his courses at the École Normale Supérieure, which I discovered while assisting his teaching and have relied on ever since. His measure-theoretic treatment of conditional expectation underlies the Conditional Expectation and MMSE Estimation chapter and informs parts of the Measure Theory and Lebesgue Integral chapters; many exercises throughout the book, especially on conditional expectation, stem from the problem sheets prepared for those courses.
I gratefully acknowledge the classical treatise of Ramis, Deschamps, and Odoux [44], which was foundational to my own learning of the Riemann integral; many results, proofs, and exercises in the Integration of Functions of a Real Variable chapter were adapted from that Riemann setting to the broader Lebesgue framework.
Finally, I am grateful to the many other authors whose foundational works on measure and integration, functional analysis and Lp spaces, and probability theory have shaped the rigor and exposition of this book [6, 8, 9, 12, 18, 19, 22, 24, 25, 26, 27, 28, 29, 33, 37, 48].
About These Topics
These graduate-level lectures on measure theory, probability, and stochastic processes are designed for students and researchers seeking a rigorous mathematical foundation for advanced applications in statistics, machine learning, information theory, queueing theory, and stochastic geometry. The material adopts a unified measure-theoretic approach, where probability theory is naturally derived from measure theory, providing a coherent framework across the three parts.
Part I covers the foundations of measure theory, the construction of the Lebesgue integral, key results such as Fubini’s theorem, the Radon-Nikodym theorem, and Lp spaces, and the classical theory of integration on the real line.
Part II develops probability theory from its measure-theoretic foundations, covering random variables and their transforms (probability generating functions, moment generating functions, characteristic functions, Laplace transforms), random vectors and Gaussian distributions, and the various modes of convergence of sequences of random variables including the Central Limit Theorem and Prohorov’s theorem.
Part III introduces stochastic processes, covering foundations through Kolmogorov’s extension theorem, stationarity, the stochastic integral, Gaussian processes, the Wiener process, conditional expectation and MMSE estimation, martingales, and ergodic theory including Birkhoff’s pointwise ergodic theorem.
Each lecture is accompanied by exercises with detailed solutions, designed to reinforce the theoretical material and build problem-solving skills. The material draws on classical references in measure theory and probability, presented in a self-contained format suitable for independent study, graduate coursework, or research preparation.
Last Updated on 8 juillet 2026 by Mohamed Kadhem KARRAY