Polimi · 2025/26
Model Identification and Data Analysis — Module 1
From white noise and stochastic processes through AR, MA and ARMA models, spectral analysis and optimal prediction to the least-squares and prediction-error identification of ARX and ARMAX models and their validation — the first module of the Politecnico di Milano Model Identification and Data Analysis course, rebuilt as an interactive, exam-focused study guide.
10 chapters~5 h reading 27 past-exam questions
After this course you can
- Characterise a stochastic process — decide stationarity and compute its mean, covariance function and power spectral density.
- Model a time series as an AR, MA or ARMA process and relate its parameters to its statistics through the Yule–Walker equations.
- Derive the optimal k-step-ahead predictor of an ARMA/ARMAX process by long division and compute its prediction-error variance.
- Identify a dynamical model from data — least squares for ARX, iterative prediction-error minimisation for ARMAX — and judge when the estimate is consistent.
- Validate an identified model with the whiteness (Anderson) test and choose its complexity with FPE/AIC/MDL or cross-validation.
Syllabus
-
Prerequisites: Linear Algebra, Probability & the Z-Transform
The mathematics MIDA1 speaks from its first lecture and never stops to teach — summation, set and quantifier notation; vectors, matrices, the inner product and the least-squares normal equations; expectation, variance and correlation; and the discrete-time signals and Z-transform every model in the course is written in — each tied to the chapter that spends it and the exam problem that grades it.
42 minutes reading. -
What Model Identification Is
The one problem behind the whole course — building a model of an uncertain dynamical system from data — and the vocabulary it runs on: white/grey/black-box modelling, static vs dynamical systems, prediction framed as an optimisation, and the white-noise residual that tells you a predictor is optimal.
low exam weight. 16 minutes reading. 1 past-exam question. -
Stochastic Processes
The language of uncertain signals: stochastic processes, their mean and autocovariance, weak vs strong stationarity, the Toeplitz condition that makes a covariance valid, white noise, and Wold's decomposition — everything the exam's first problem is built on.
high exam weight. 32 minutes reading. 3 past-exam questions. -
MA, AR & ARMA Models
The model zoo that turns Wold's "white noise through a filter" into a handful of parameters: moving-average (MA), auto-regressive (AR) and ARMA processes, the monic normalisation, the Yule–Walker equations, and the covariance fingerprints that tell the three apart.
medium exam weight. 30 minutes reading. 4 past-exam questions. -
Frequency Analysis & Spectral Factorization
The frequency-domain view of a stochastic process: the power spectral density and its properties, the master formula that turns a filter into a spectrum, the model-zoo spectra, why the periodogram never converges, and the spectral factorization that puts a process into the canonical form prediction needs.
high exam weight. 34 minutes reading. 2 past-exam questions. -
The Prediction Problem
The Kolmogorov–Wiener theory that is the heart of MIDA1: the optimal k-step predictor of an ARMA/ARMAX process, built by long division of C/A (equivalently the Diophantine equation), read from data through the whitening filter, with its error variance — and why the canonical form is mandatory.
high exam weight. 38 minutes reading. 4 past-exam questions. -
Identification: Least Squares & ARX
The linear half of system identification: the least-squares estimator and its normal equations, the disturbance-based taxonomy of model families, the predictive approach that turns "which model is best" into an optimisation, and the key result that an ARX predictor is a linear regression you solve in one shot.
high exam weight. 34 minutes reading. 5 past-exam questions. -
PEM: Asymptotics, ARMAX & Identifiability
The deeper half of identification: what prediction-error minimisation converges to as data grows (consistency when the true system is in the model class, the best approximant otherwise), experimental vs structural identifiability, the uncertainty of the estimates, and the iterative maximum-likelihood algorithm that ARMAX's nonlinear predictor forces on us.
high exam weight. 38 minutes reading. 4 past-exam questions. -
Model Validation & Selection
Closing the identification loop: is the model good, and how complex should it be? The whiteness (Anderson) test on the residual with its confidence band, the cross-correlation test for the input path, and the FPE / AIC / MDL criteria and cross-validation for choosing the model order.
medium exam weight. 26 minutes reading. 3 past-exam questions. -
Time-Series Analysis & Practical Aspects
The dedicated time-series toolkit and the engineering that makes identification work in practice: Yule–Walker and the Durbin–Levinson recursion, the PARCOR function that reads off AR order, differencing a non-stationary series into an ARIMA model, and designing an informative experiment (input richness, sampling time, pre-filtering).
low exam weight. 24 minutes reading. 1 past-exam question.