templates
3.1 matrix analyses
There is one matrix-analysis function in this course: runMatrixAnalysis(). If you meet older
material calling runMatrixAnalyses(), with an “es” on the end, that was a wide-data-only version and
it was retired in September 2026 — calling it now returns a message telling you what to use instead.
3.1.1 basic template: a wide data frame
A wide data frame has one row per sample and one column per analyte. chemical_blooms is exactly
that shape: a label column naming each plant, and then one column for each compound class.
runMatrixAnalysis(
data = chemical_blooms,
analysis = "pca",
columns_w_values_for_single_analyte = c(
"Alkanes", "Sec_Alcohols", "Others", "Fatty_acids", "Alcohols",
"Triterpenoids", "Ketones", "Other_compounds", "Aldehydes"
),
columns_w_sample_ID_info = c("label")
)That is the whole basic call, and it is only four questions:
-
data— which table. -
analysis— which analysis."pca","hclust","kmeans","dist"and the rest are listed in the full template below. -
columns_w_values_for_single_analyte— which columns hold the numbers. One column per analyte. -
columns_w_sample_ID_info— which column (or columns) name a sample. If it takes two columns together to identify a sample, list both.
If your data are long instead — one row per sample per analyte, with the analyte names stacked in
a single column — then swap columns_w_values_for_single_analyte for the pair
column_w_names_of_multiple_analytes and column_w_values_for_multiple_analytes. Everything else
stays the same.
3.1.2 full template
runMatrixAnalysis(
data, # data to use for analysis
analysis = c(
"pca", "pca_ord", "pca_dim", # PCA
"mca", "mca_ord", "mca_dim", # MCA (PCA on categorical data)
"mds", "mds_ord", "mds_dim", # MDS
"tsne", "dbscan", "kmeans", # Clustering
"hclust", "hclust_phylo" # Hierarchical clustering
),
parameters = NULL,
column_w_names_of_multiple_analytes = NULL,
column_w_values_for_multiple_analytes = NULL,
columns_w_values_for_single_analyte = NULL,
columns_w_additional_analyte_info = NULL,
columns_w_sample_ID_info = NULL,
transpose = FALSE, # default = FALSE, this chooses whether to transpose the data
distance_method = c( # the distance metric to use in computing a distance matrix
"euclidean", "manhattan",
"gower" # for mixed numeric / categorical data; computed with cluster::daisy
),
agglomeration_method = c( # the clustering method to use in heirarchical clustering
"ward.D2", "ward.D", "single", "complete",
"average", # (= UPGMA)
"mcquitty", # (= WPGMA)
"median", # (= WPGMC)
"centroid" # (= UPGMC)
),
tree_method = c( # how a tree is built, for analysis = "hclust" / "hclust_phylo"
"linkage_dendrogram", # DEFAULT. Hierarchical clustering proper: uses agglomeration_method,
# and returns bootstrap support values.
"neighbor_joining" # The phylogenetic method. Works from the distance matrix alone, so it
# IGNORES agglomeration_method, and gives no bootstrap values.
),
unknown_sample_ID_info = NULL,
components_to_return = 2, # how many principal components to return
scale_variance = NULL, # default = TRUE for EVERY analysis: centers each column and divides it by
# its standard deviation, so no single large-valued variable dominates the
# distances. Pass FALSE for compositional data (every column a share of the
# same whole), where the columns are already on one scale.
na_replacement = c("mean", "none", "zero", "drop"), # default = "mean", this chooses what to do with missing values
output_format = c("wide", "long"), # default = "wide". For analysis = "dist" this chooses the
# RETURN TYPE: "wide" gives a base R `dist` object, "long"
# gives one row per PAIR of samples, which is what you want
# if you are about to build a network from the distances.
)