Linear Regression with Gradient Descent: Mathematical Foundations of AI
Preface
The Book
Why bother with linear regression, a model that has been around for more than a century? After all, ChatGPT or Claude are not linear regressions. They are large Transformer-based neural networks, also called Large Language Models (LLMs).
You may be able to understand such models without knowing linear regression. But you may find yourself struggling when reading terms like cross-entropy or gradient descent.
By developing an intuitive understanding of linear regression, we will cover the concepts at the core of the current AI revolution:
- Model fitting and evaluation
- Functions and derivatives
- The chain rule
- Gradient descent
- Logistic regression
- and a lot more…
How to read this book?
It is written as a story with chapters building on top of one another. The maths covered should be relatively self-contained. If the beginning feels a bit too easy, feel free to skip ahead.
About Me
I am a Machine Learning Scientist and Computer Science Lecturer in Berlin. I have a blog and newsletter, subscribe here: https://eliottkalfon.com to read my latest work.
Notes
- All the data in this book is synthetic and generated by me
- If you find yourself struggling with a chapter or formula, I recommend taking a break and doing some research
- Also, if this happens too many times, please reach out at https://eliottkalfon.com/contact