Linear Algebra and Learning from Data
The linear algebra that deep learning is actually built on.
What is actually freeMIT does not host the full book. What is published free is the Table of Contents, the Preface, the Deep Learning essay, three complete sample sections (I.1, I.2 and VII.1), the errata, and three sets of solutions. Fully free, however, is MIT 18.065, the course Strang taught directly from this book: 36 video lectures plus problem sets, mapped section by section below. So for most sections the video and problems are your free route; a full section PDF exists only for the three samples. Everything here links to MIT's own servers. If you can afford the book, buy it. That supports the author.
- Full chapter, free to read 3
- Covered by a lecture 30
- Problem set and contents only 13
Most of this book is followed through the lectures rather than the page. That is the honest shape of what MIT publishes, and it is why the roadmap is built around the course rather than around chapter downloads.
Highlights of Linear Algebra
Column spaces, factorizations, eigenvalues, SVD: the foundation.
Part I progress: 0/11 sections read
The Four Fundamental Subspaces
Elimination and A = LU
Orthogonal Matrices and Subspaces
Eigenvalues and Eigenvectors
Symmetric Positive Definite Matrices
Rayleigh Quotients and Generalized Eigenvalues
Norms of Vectors and Functions and Matrices
Before moving on from Part I
Computations with Large Matrices
Numerical linear algebra, least squares, randomized methods.
Part II progress: 0/4 sections read
Three Bases for the Column Space
Randomized Linear Algebra
Before moving on from Part II
Low Rank and Compressed Sensing
How matrices change, interlacing eigenvalues, decaying singular values.
Part III progress: 0/5 sections read
Rapidly Decaying Singular Values
Split Algorithms for ℓ² + ℓ¹
Compressed Sensing and Matrix Completion
Before moving on from Part III
Special Matrices
Circulants, Fourier, graphs, clustering, distance matrices.
Part IV progress: 0/10 sections read
Fourier Transforms: Discrete and Continuous
The Kronecker Product A ⊗ B
Sine and Cosine Transforms from Kronecker Sums
Toeplitz Matrices and Shift Invariant Filters
Graphs and Laplacians and Kirchhoff's Laws
Clustering by Spectral Methods and k-means
Completing Rank One Matrices
The Orthogonal Procrustes Problem
Before moving on from Part IV
Probability and Statistics
Mean, variance, covariance: the statistical toolkit behind learning.
Part V progress: 0/6 sections read
Mean, Variance, and Probability
Probability Distributions
Moments, Cumulants, and Inequalities of Statistics
Covariance Matrices and Joint Probabilities
Multivariate Gaussian and Weighted Least Squares
Before moving on from Part V
Optimization
Convexity, Lagrange multipliers, gradient descent, duality.
Part VI progress: 0/5 sections read
Minimum Problems: Convexity and Newton's Method
Lagrange Multipliers = Derivatives of the Cost
Linear Programming, Game Theory, and Duality
Stochastic Gradient Descent and ADAM
Before moving on from Part VI
Learning from Data
Neural network architecture, convolutions, backpropagation.
Part VII progress: 0/5 sections read
Backpropagation and the Chain Rule
Hyperparameters: The Fateful Decisions
The World of Machine Learning
Before moving on from Part VII
Material for the whole book
Table of Contents, from Gilbert Strang / MIT, opens in a new tab
Preface, from Gilbert Strang / MIT, opens in a new tab
Deep Learning and Neural Nets, from Gilbert Strang / MIT, opens in a new tab
Functions of Deep Learning (SIAM), from Gilbert Strang / SIAM News, opens in a new tab
Solutions by Strang, from Gilbert Strang / MIT, opens in a new tab
Solutions (Chang & Davidson), from Chang & Davidson / MIT, opens in a new tab
Solutions (Tohme), from Tohme / MIT, opens in a new tab
All problem sets (one PDF), from MIT OpenCourseWare, opens in a new tab
All 36 video lectures, from MIT OpenCourseWare, opens in a new tab
YouTube playlist, from MIT OpenCourseWare, opens in a new tab
OCW reading map, from MIT OpenCourseWare, opens in a new tab
Counting Parameters, from Gilbert Strang / MIT, opens in a new tab
Central Limit Theorem (p.288), from Gilbert Strang / MIT, opens in a new tab
Convolution as a Moving Window, from Gilbert Strang / MIT, opens in a new tab
Errata, from Gilbert Strang / MIT, opens in a new tab
Original MIT book page, from Gilbert Strang / MIT, opens in a new tab
Attribution
Linear Algebra and Learning from Data is copyright 2019 Gilbert Strang, published by Wellesley-Cambridge Press. This page is an independent reading guide. It reproduces no part of the book and hosts no files: every link above points to the author's own page or to the publishing university's servers.
Lecture and problem links go to MIT 18.065, Matrix Methods in Data Analysis, Signal Processing, and Machine Learning (Spring 2018), used under CC BY-NC-SA 4.0.
If this book is useful to you and you can afford it, buy a copy. That is what keeps authors writing them.