19/05/2026
đPrincipal Component Analysisđ
Principal Component Analysis (PCA) is a powerful dimensionality reduction technique used to simplify complex datasets by transforming many correlated variables into a smaller number of uncorrelated variables called principal components. PCA is widely used in agriculture, genomics, bioinformatics, image analysis, and machine learning.
đ What is PCA?
PCA converts a dataset with many variables into a new coordinate system where:
PC1 (Principal Component 1) explains the largest amount of variation.
PC2 explains the second largest amount of variation.
PC3 explains the next largest variation.
And so on.
These components are linear combinations of the original variables.
đ Why PCA is Important
PCA is useful for:
Reducing dimensionality
Removing redundancy among correlated variables
Visualizing high-dimensional data
Detecting outliers
Identifying variable relationships
Preprocessing data before machine learning
đ Applications of PCA
Agriculture
Soil property analysis
Nutrient profiling
Plant phenotyping
Spectral data analysis
Genomics and Bioinformatics
Gene expression analysis
Population genetics
SNP analysis
Image Processing
Feature extraction
Compression
Finance
Stock market analysis
Environmental Science
Climate and pollution studies
đ Example Dataset
Suppose you have plant measurements:
Plant Height LeafArea Chlorophyll RootLength
1 35 120 42 15
2 40 135 46 18
3 38 128 44 17
4 45 150 49 21
đ Install Required Packages
pip install pandas numpy matplotlib scikit-learn
đ Python Code for PCA
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
# Load dataset
df = pd.read_csv("plant_data.csv")
# Optional: separate treatment labels if present
# labels = df['Treatment']
# X = df.drop('Treatment', axis=1)
X = df # if all columns are numeric
# Standardize the data
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
# Perform PCA
pca = PCA()
X_pca = pca.fit_transform(X_scaled)
# Explained variance ratio
explained_variance = pca.explained_variance_ratio_
print("Explained variance ratio:")
print(explained_variance)
# Create scores dataframe
scores = pd.DataFrame(
X_pca,
columns=[f"PC{i+1}" for i in range(X_pca.shape[1])]
)
print(scores.head())
đ Scree Plot
plt.figure(figsize=(8, 5))
plt.plot(
range(1, len(explained_variance) + 1),
explained_variance,
marker='o'
)
plt.xlabel("Principal Component")
plt.ylabel("Explained Variance Ratio")
plt.title("Scree Plot")
plt.show()
đ PCA Score Plot
plt.figure(figsize=(8, 6))
plt.scatter(scores['PC1'], scores['PC2'])
for i in range(len(scores)):
plt.text(scores['PC1'][i], scores['PC2'][i], str(i+1))
plt.xlabel("PC1")
plt.ylabel("PC2")
plt.title("PCA Score Plot")
plt.grid(True)
plt.show()
đ PCA with Treatment Groups
labels = df['Treatment']
X = df.drop('Treatment', axis=1)
X_scaled = StandardScaler().fit_transform(X)
pca = PCA(n_components=2)
X_pca = pca.fit_transform(X_scaled)
scores = pd.DataFrame(X_pca, columns=['PC1', 'PC2'])
scores['Treatment'] = labels
for treatment in scores['Treatment'].unique():
subset = scores[scores['Treatment'] == treatment]
plt.scatter(subset['PC1'], subset['PC2'], label=treatment)
plt.xlabel("PC1")
plt.ylabel("PC2")
plt.legend()
plt.title("PCA by Treatment")
plt.show()
đ Loadings (Variable Contributions)
loadings = pd.DataFrame(
pca.components_.T,
columns=['PC1', 'PC2'],
index=X.columns
)
print(loadings)
Interpretation:
Large positive or negative values indicate strong influence.
Variables with similar directions are positively correlated.
Opposite directions indicate negative correlation.
đ Biplot
plt.figure(figsize=(8, 8))
plt.scatter(scores['PC1'], scores['PC2'])
for i, var in enumerate(X.columns):
plt.arrow(
0, 0,
loadings.iloc[i, 0] * 3,
loadings.iloc[i, 1] * 3,
head_width=0.05
)
plt.text(
loadings.iloc[i, 0] * 3.2,
loadings.iloc[i, 1] * 3.2,
var
)
plt.xlabel("PC1")
plt.ylabel("PC2")
plt.title("PCA Biplot")
plt.grid(True)
plt.show()
đ Selecting the Number of Components
A common rule is to retain enough components to explain 80â95% of the total variance.
import numpy as np
cumulative_variance = np.cumsum(explained_variance)
print(cumulative_variance)
đ Interpretation Example
Suppose:
PC1 explains 65%
PC2 explains 20%
Total = 85%
This means the first two components summarize most of the information in the dataset.
đ PCA in Agriculture and Bioinformatics
For your research, PCA can be used to analyze:
Soil nutrient concentrations (N, P, K)
Moisture and fertigation responses
Spectral signatures
Gene expression profiles
Plant phenotyping traits
Root architecture measurements
đ Advantages of PCA
Reduces complexity
Removes multicollinearity
Improves visualization
Speeds up machine learning algorithms
đ Limitations of PCA
Assumes linear relationships
Components may be harder to interpret biologically
Sensitive to scaling
Can be affected by outliers