OPTIMIZATION METHODS FOR DATA SCIENCE Canale unico

Docente coordinatore e verbalizzante: SAVERIO SALZO

Docenti

Obiettivi formativi

General goals: The aim of the course is to introduce students to the theory and applications of optimization techniques for machine learning problems. Also, students are expected to acquire knowledge about standard models used in machine learning, such as Deep Neural Networks and Support Vector Machines.

Specific goals: The course will put a special emphasis on convex optimization techniques which play a key role in data sciences. The objective of this course is to provide basic tools and methods at the core of modern nonlinear convex optimization. Starting from the gradient descent method we will cover some state of the art algorithms, including proximal gradient methods, accelerated methods, stochastic subgradient method and randomized block-coordinate descent methods, which are nowadays very popular techniques to solve machine learning and inverse problems. The course will also cover topics in statistics in high dimension which is the typical scenario in machine learning problems.

Knowledge and understanding: the student will learn about (1) mathematical formulation of the machine learning problem, (2) the latest optimization algorithms for solving machine learning problems and extract information from data (3) understanding the theory of convergence of such algorithms (4) develop computational skills for handling several important problems in data science.

Applying knowledge and understanding: through several examples from applied sciences and lab sessions, the student will appreciate the importance of optimization techniques and will understand which algorithm is most appropriate to use in each context.

Critical and judgmental skills: the student will be able to tackle with rigor a number of significant optimization problems and algorithms so as to become fully aware of the technicalities and main ideas behind the various approaches. This will stimulate the student's independent judgment.

Communication skills: by studying the theoretical and practical aspects of optimization techniques the student will learn gradually to communicate with rigor and clarity. She will also learn that a proper understanding of the mathematical aspects of data science is one of the main skills to achieve effective communication.

Learning skills: students will have the chance to have additional details on some specific topics.

Risultati di apprendimento attesi

Knowledge and understanding: students will acquire an advanced understanding of mathematical and algorithmic foundations underlying optimization techniques used in data science. They will:
- understand key concepts in convex optimization (smooth and nonsmooth), constrained and unconstrained optimization, and stochastic methods (marginally, non-convex optimization will be covered too).
- comprehend how optimization principles relate to machine learning models, statistical estimation, and data-driven decision processes.
- gain familiarity with computational complexity issues, convergence guarantees, and practical algorithm implementation aspects (e.g., gradient descent, coordinate methods, proximal algorithms).

Applying knowledge and understanding: students will be able to:
- formulate real-world data science problems as optimization problems and select suitable algorithmic strategies to solve them.
- implement and evaluate optimization algorithms in computational environments (e.g., MATLAB).
- apply optimization tools to tasks such as regression, classification, and regularized approaches.

Making judgements: students will develop the ability to:
- critically analyze the suitability and limitations of different optimization methods for specific data-driven tasks.
- evaluate trade-offs between computational efficiency, convergence properties, and model accuracy.
- interpret optimization results within a data science workflow and make informed methodological choices.
- identify when heuristic, approximate, or hybrid optimization approaches are appropriate.

Communication skills: students will be able to:
- clearly communicate optimization principles, assumptions, and results to both technical and non-technical audiences.
- prepare concise reports and visualizations describing algorithmic design choices and experimental findings.

Learning skills: students will develop the capacity to:
- independently explore and learn emerging optimization frameworks and software relevant to data science.
- extend foundational knowledge to advanced topics such as distributed optimization and large-scale learning.
- critically engage with scientific literature and ongoing research in optimization and data-driven modeling.

Prerequisiti

Analisi per funzioni di più variabili, algebra lineare, probabilità

Programma dell’insegnamento

1. Basics on convex sets and functions.
2. Differentiability and convexity
3. Smooth convex optimization
3. Differential theory for nonsmooth convex functions
4. Duality theory (part 1).
5. The projected subgradient method.
6. Frank-Wolfe algorithm
7. Proximity operators and averaged operators
8. The proximal gradient algorithm
9. Accelerated proximal gradient algorithms
10. Elements of sparse estimation/recovery
11. Proximal methods for convex spectral functions
12. Stochastic optimization algorithms: SGD and randomized methods
13 Duality Theory (part 2)
14 Dual algorithms
15. Applications in machine learning: Support vector machine and concentration inequalities

Testi di riferimento

S. Boyd and L. Vandenberghe, Convex Optimization, Cambridge University Press, 2004

Y. Nesterov, Introductory Lectures on Convex Optimization, Kluwer, 2024

S. Salzo, S. Villa, Proximal Gradient Methods for Machine Learning and Imaging, 2022

Lecture notes by Prof. Salzo

Modalità di svolgimento

Lectures will be only in presence.

Frequenza

Nessun obbligo di frequenza

Modalità di esame

Continuous Assessment (40%)
Continuous assessment monitors students’ progress throughout the course and may include:
- homework assignments / mini-projects
- implementation and analysis of optimization algorithms (e.g., gradient descent, Newton’s method, stochastic optimization) on real or synthetic data.
- short written reports or notebooks documenting formulation, results, and discussion of findings.
- in-class exercises or quizzes.

Written/Oral Exam (60%):
- a 2-hour written exam including both theoretical questions (proof-based or conceptual) and applied problems (algorithmic exercises, short computations).
- optional oral discussion to verify understanding and reasoning depth.

Esempi di domande

Prove that the function f (x, y) = |x − 2y|^22 is (jointly) convex.

Programmazione delle attività didattiche

  • Introduction and motivations.  Preliminaries: calculus in several variables, probability and linear algebra (Week 1)
    • Testi di riferimento: Lecture Notes from the Instructor

  • Convex sets and functions: main properties. (Week 2)
    • Testi di riferimento: Lecture Notes from the Instructor

  • Smooth optimization: the gradient descent algorithm and its analysis of convergence (Week 3)
    • Testi di riferimento: Lecture Notes from the Instructor

  • Nonsmooth differential theory (Week 4)
    • Testi di riferimento: Lecture Notes from the Instructor

  • Duality Theory (part 1): The Fenchel transform and its properties. (Week 5)
    • Testi di riferimento: Lecture Notes from the Instructor

  • Nonsmooth optimization methods: proximal gradient methods and Frank-Wolfe Algorithm (Week 6)
    • Testi di riferimento: Lecture Notes from the Instructor

  • Sparsity and the statistics of Lasso. Convex functions of matrices (Week 7)
    • Testi di riferimento: Lecture Notes from the Instructor

  • Stochastic Optimization algorithms and randomized methods. (Week 8)
    • Testi di riferimento: Lecture Notes from the Instructor

  • Duality Theory (part 2). The Fenchel-Rochafellar duality and dual algorithms (Week 9)
    • Testi di riferimento: Lecture Notes from the Instructor

  • Application in Machine Learning and Image Processing: Concentration Inequalities, Empirical Risk Minimization, Kernel Methods and SVM's (Week 10)
    • Testi di riferimento: Lecture Notes from the Instructor

Obiettivi per lo sviluppo sostenibile - Agenda ONU 2030

  • Goal5
  • Goal9
  • Goal10
  • Anno accademico2026/2027
  • Corso di studio a cui afferisce l’insegnamentoData Science
  • Codice insegnamento10606725
  • Anno e semestre1º anno - 2º semestre
  • TipologiaAttività formative caratterizzanti
  • AmbitoFormazione matematico-statistica
  • SSDMAT/09
  • Presenza obbligatoriaNo
  • Linguaeng
  • CFU6 CFU
  • Durata complessiva60 ore
  • Distribuzione delle ore36 classroom hours, 24 training hours