Lecture 4: More PCA, LSA, and NMF#

UBC Master of Data Science program, 2024-25

Imports and learning outcomes#

Imports#

import os
import random
import sys
import time

import numpy as np

sys.path.append(os.path.join(os.path.abspath(".."), "code"))
import matplotlib.pyplot as plt
import seaborn as sns
import torch
from plotting_functions import *
from sklearn.decomposition import PCA
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

plt.rcParams["font.size"] = 16
# plt.style.use("seaborn")
%matplotlib inline
pd.set_option("display.max_colwidth", 0)

DATA_DIR = os.path.join(os.path.abspath(".."), "data/")
---------------------------------------------------------------------------
ImportError                               Traceback (most recent call last)
Cell In[1], line 11
      9 import matplotlib.pyplot as plt
     10 import seaborn as sns
---> 11 import torch
     12 from plotting_functions import *
     13 from sklearn.decomposition import PCA

File ~/miniforge3/envs/jbook/lib/python3.12/site-packages/torch/__init__.py:367
    365     if USE_GLOBAL_DEPS:
    366         _load_global_deps()
--> 367     from torch._C import *  # noqa: F403
    370 class SymInt:
    371     """
    372     Like an int (including magic methods), but redirects all operations on the
    373     wrapped node. This is used in particular to symbolically record operations
    374     in the symbolic shape workflow.
    375     """

ImportError: dlopen(/Users/kvarada/miniforge3/envs/jbook/lib/python3.12/site-packages/torch/_C.cpython-312-darwin.so, 0x0002): Library not loaded: @rpath/libgfortran.5.dylib
  Referenced from: <0B9C315B-A1DD-3527-88DB-4B90531D343F> /Users/kvarada/miniforge3/envs/jbook/lib/libopenblas.0.dylib
  Reason: tried: '/Users/kvarada/miniforge3/envs/jbook/lib/libgfortran.5.dylib' (duplicate LC_RPATH '@loader_path'), '/Users/kvarada/miniforge3/envs/jbook/lib/libgfortran.5.dylib' (duplicate LC_RPATH '@loader_path'), '/Users/kvarada/miniforge3/envs/jbook/lib/libgfortran.5.dylib' (duplicate LC_RPATH '@loader_path'), '/Users/kvarada/miniforge3/envs/jbook/lib/libgfortran.5.dylib' (duplicate LC_RPATH '@loader_path'), '/Users/kvarada/miniforge3/envs/jbook/lib/libgfortran.5.dylib' (duplicate LC_RPATH '@loader_path'), '/Users/kvarada/miniforge3/envs/jbook/lib/libgfortran.5.dylib' (duplicate LC_RPATH '@loader_path'), '/Users/kvarada/miniforge3/envs/jbook/lib/libgfortran.5.dylib' (duplicate LC_RPATH '@loader_path'), '/Users/kvarada/miniforge3/envs/jbook/lib/libgfortran.5.dylib' (duplicate LC_RPATH '@loader_path'), '/Users/kvarada/miniforge3/envs/jbook/lib/python3.12/site-packages/torch/lib/libgfortran.5.dylib' (no such file), '/Users/kvarada/miniforge3/envs/jbook/lib/python3.12/site-packages/torch/../../../libgfortran.5.dylib' (duplicate LC_RPATH '@loader_path'), '/Users/kvarada/miniforge3/envs/jbook/lib/python3.12/site-packages/torch/lib/libgfortran.5.dylib' (no such file), '/Users/kvarada/miniforge3/envs/jbook/lib/python3.12/site-packages/torch/../../../libgfortran.5.dylib' (duplicate LC_RPATH '@loader_path'), '/Users/kvarada/miniforge3/envs/jbook/bin/../lib/libgfortran.5.dylib' (duplicate LC_RPATH '@loader_path'), '/Users/kvarada/miniforge3/envs/jbook/bin/../lib/libgfortran.5.dylib' (duplicate LC_RPATH '@loader_path'), '/usr/local/lib/libgfortran.5.dylib' (no such file), '/usr/lib/libgfortran.5.dylib' (no such file, not in dyld cache)

Learning outcomes#

From this lecture, students are expected to be able to:

  • Explain how to write a training example as a linear combination of components, in the context of PCA;

  • Use explained_variance_ratio_ of PCA to choose \(k\) in the context of dimensionality reduction using PCA.

  • Explain the geometry of PCA for different values for \(n\) and \(d\).

  • Explain similarities and differences between K-Means and PCA.

  • State similarities and differences between PCA, TruncatedSVD, and NMF.

  • Interpret components givens given by PCA, Truncated SVD and NMF.

  • Recognize the scenarios for using TruncatedSVD and NMF and use them with scikit-learn.



1. PCA recap#

A commonly used dimensionality reduction technique.

  • PCA input

    • \(X_{n \times d}\)

    • the number of components \(k\)

  • PCA output

    • \(W_{k \times d}\) with \(k\) principal components (also called “parts”)

    • \(Z_{n \times k}\) transformed data (also called “part weights”)

  • Usually \(k << d\).

  • PCA is useful for dimensionality reduction, visualization, feature extraction etc.

Important matrices involved in PCA

PCA learns a \(k\)-dimensional subspace of the original \(d\)-dimensional space. Here are the main matrices involved in PCA.

  • \(X \rightarrow\) Original data matrix

  • \(W \rightarrow\) Principal components

    • the best lower-dimensional hyper-plane found by the PCA algorithm

  • \(Z \rightarrow\) Transformed data with reduced dimensionality

    • the co-ordinates in the lower dimensional space

  • \(X_{hat} \rightarrow\) Reconstructed data

    • reconstructions using \(Z\) and \(W\) in the original co-ordinate system

How can we use \(W\) and \(Z\)?

  • Dimension reduction: compress data into limited number of components.

  • Outlier detection: it might be an outlier if isn’t a combination of usual components.

  • Supervised learning: we could use \(Z\) as our features.

  • Visualization: if we have only 2 components, we can view data as a scatterplot.

  • Interpretation: we can try and figure out what the components represent.

  • PCA reduces the dimensionality by learning a \(k\)-dimensional subspace of the original \(d\)-dimensional space.

  • When going from higher dimensional space to lower dimensional space, PCA still tries to capture the topology of the points in high dimensional space, making sure that we are not losing some of the important properties of the data.

  • So Points which are nearby in high dimensions are still nearby in low dimension.

  • In PCA, we find a lower dimensional subspace so that the squared error of reconstruction, i.e., elements of \(X\) and elements of \(ZW\) is minimized.

  • The goal is to find the two best matrices such that when we multiply them we get a matrix that’s closest to the data.

  • A common way to learn PCA is using singular value decomposition (SVD).

PCA examples

  • Let’s look at examples of reducing dimensionality using PCA from 3 to 2 and from 3 to 1.

n = 12
d = 3

x1 = np.linspace(0, 5, n) + np.random.randn(n) * 0.05
x2 = -x1 * 0.1 + np.random.randn(n) * 2
x3 = x1 * 0.7 + np.random.randn(n) * 3

X = np.concatenate((x1[:, None], x2[:, None], x3[:, None]), axis=1)
X = X - np.mean(X, axis=0)
import plotly.express as px

plot_interactive_3d(X)