Contextualized sensorimotor norms: Multi-dimensional measures of sensorimotor strength for ambiguous English words, in context.
The 6 matches
- [1] § Case study: Can large language models help scale the CS norms? › Between-dimension comparison ↔ src/analysis/contextualized_norms_analysis.Rmd, lines 2313–2381 · score 0.67 · Mantel, permutations, pairwise, Pearson, matrices, Spearman
- [2] § Contextual variance in sensorimotor associations › Sense dominance and deviation from the Lancaster norms ↔ src/analysis/contextualized_norms_analysis.Rmd, lines 1170–1200 · score 0.66 · decontextualized LS norm, cosine distance, sensorimotor properties, lme4, dominant sense, closer
- [3] § Comparing distributional and sensorimotor similarity › Methods › Calculating distributional distance ↔ src/analysis/contextualized_norms_analysis.Rmd, lines 1254–1326 · score 0.65 · DistilBERT, RoBERTa, ALBERT, RAW, distance, models
- [4] § Contextual variance in sensorimotor associations › Deviations from the Lancaster sensorimotor norms ↔ src/analysis/contextualized_norms_analysis.Rmd, lines 766–891 · score 0.64 · standardized absolute deviation, maximally, floor, Olfaction, SD, Taste
- [5] § Comparing distributional and sensorimotor similarity › Sensorimotor distance and ambiguity processing › Explanation of experimental materials ↔ src/analysis/contextualized_norms_analysis.Rmd, lines 1870–1918 · score 0.62 · primed sensibility judgment, Log RT, aggregated, accuracy, asked, word
- [6] § Comparing distributional and sensorimotor similarity › Sensorimotor distance and ambiguity processing › Evaluating sensorimotor vs. distributional distance ↔ src/analysis/contextualized_norms_analysis.Rmd, lines 1870–1918 · score 0.56 · primed sensibility judgment, Log RT, accuracy, sensorimotor distance
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
R Markdown · 2,607 lines · 81 KB · no license · 6 matches
- ---
- title: "Analysis of contextualized sensorimotor norms"
- author: "Sean Trott"
- date: "January 19, 2026"
- output:
- html_document:
- keep_md: yes
- toc: yes
- toc_float: yes
- ---
- ```{r setup, include=FALSE}
- knitr::opts_chunk$set(dpi = 300, fig.format = "png")
- ```
- ```{r include=FALSE}
- library(tidyverse)
- library(lme4)
- library(ggridges)
- library(broom.mixed)
- library(lmerTest)
- library(ggcorrplot)
- library(kableExtra)
- ```
- # Introduction
- In this document, we analyze the **contextualized sensorimotor norms**: judgments about the strength of different sensorimotor dimensions of ambiguous words, in context.
- We use these norms in several analyses:
- 1. First, we compare them to the corresponding dimensions for the "decontextualized" Lancaster sensorimotor norms.
- 2. Second, we ask whether the **dominance** of a word sense is correlated with its sensorimotor strength, i.e., whether more concrete meanings tend to be rated as more dominant.
- 3. Third, we ask whether the **sensorimotor distance** between two contexts of use predicts judgments of how **related** those meanings are, above and beyond their distributional similarity and whether or not they belong to the same sense.
- 4. Fourth, we use **sensorimotor distance** to predict behavior on a primed sensibility judgment task.
- # Characterizing dimensions
- First, load the data.
- ```{r}
- # setwd("/Users/seantrott/Dropbox/UCSD/Research/Ambiguity/SSD/cs_norms/src/analysis")
- df_contextualized_meanings = read_csv("../../data/processed/contextualized_sensorimotor_norms.csv")
- nrow(df_contextualized_meanings)
- ```
- ## Visualizing distributions
- ```{r different_dimensions}
- df_contextualized_meanings_long = df_contextualized_meanings %>%
- pivot_longer(cols = c(Vision.M, Hearing.M, Olfaction.M,
- Taste.M, Interoception.M, Touch.M,
- Mouth_throat.M, Head.M, Torso.M,
- Hand_arm.M, Foot_leg.M),
- names_to = "Dimension",
- values_to = "Strength") %>%
- mutate(Dimension = gsub('.M', '', Dimension))
- df_contextualized_meanings_long %>%
- ggplot(aes(x = reorder(Dimension, Strength),
- y = Strength)) +
- geom_violin() +
- geom_jitter(alpha = .1,
- width = .1) +
- coord_flip() +
- labs(y = "Sensorimotor strength",
- x = "Dimension") +
- theme_bw() +
- theme(text = element_text(size=20))
- ```
- ## Correlations across dimensions
- ```{r corr_matrix}
- columns = df_contextualized_meanings %>%
- mutate(Vision = Vision.M,
- Hearing = Hearing.M,
- Olfaction = Olfaction.M,
- Taste = Taste.M,
- Interoception = Interoception.M,
- Touch = Touch.M,
- Mouth_throat = Mouth_throat.M,
- Head = Head.M,
- Torso = Torso.M,
- Hand_arm = Hand_arm.M,
- Foot_leg = Foot_leg.M) %>%
- select(Vision, Hearing, Olfaction,
- Taste, Interoception, Touch,
- Mouth_throat, Head, Torso,
- Hand_arm, Foot_leg)
- cors = cor(columns)
- # cors[lower.tri(cors, diag=TRUE)] <- 0
- # Plot the correlation matrix
- ggcorrplot(cors,
- hc.order = FALSE,
- # method = "square",
- type = "upper") +
- theme(
- axis.text.x = element_text(size = 10, angle = 45, hjust = 1),
- axis.text.y = element_text(size = 10)
- )
- ggcorrplot(cors, hc.order = FALSE, type = "upper",
- lab = TRUE, lab_size = 2.5) +
- theme(axis.text.x = element_text(size = 10, angle = 45, hjust = 1))
- ```
- # Predicting dominance
- ## Load and merge data
- Load the item-level means for the sensorimotor norms.
- ```{r}
- df_contextualized_meanings = read_csv("../../data/processed/contextualized_sensorimotor_norms_with_ls.csv")
- nrow(df_contextualized_meanings)
- ```
- Load the dominance norms.
- ```{r}
- df_dominance = read_csv("../../data/processed/dominance_norms_with_order.csv")
- ## Determine the specific sense/meaning of the righthand context
- df_dominance = df_dominance %>%
- mutate(context = substr(version_with_order, 6, 9))
- ## Now group by that righthand context to get relative dominance of that meaning
- df_dominance_individual = df_dominance %>%
- group_by(word, context) %>%
- summarise(dominance = mean(dominance_right))
- nrow(df_dominance_individual)
- ```
- Merge the dominance and sensorimotor norms data.
- ```{r}
- df_dom_plus_sm = df_contextualized_meanings %>%
- inner_join(df_dominance_individual)
- nrow(df_dom_plus_sm)
- ```
- We also load and merge the Lancaster norms, as a control.
- ```{r}
- df_lancaster <- read_csv("../../data/lexical/lancaster_norms.csv") %>%
- mutate(word = tolower(Word)) %>%
- rename(
- Foot_leg.SD.LS = Foot_leg.SD,
- Hand_arm.SD.LS = Hand_arm.SD,
- Head.SD.LS = Head.SD,
- Torso.SD.LS = Torso.SD
- )
- df_dom_plus_sm <- df_dom_plus_sm %>%
- inner_join(df_lancaster)
- nrow(df_dom_plus_sm)
- ```
- ## Calculating contextualized sensorimotor strength
- Based on Lynott et al (2019), we **contextualized sensorimotor strength** as the *maximum* strength across all the dimensions.
- ```{r sm_strength_overall}
- df_dom_plus_sm = df_dom_plus_sm %>%
- rowwise() %>%
- mutate(max_strength = max(
- c(
- ## Modalities
- Vision.M,
- Hearing.M,
- Olfaction.M,
- Touch.M,
- Taste.M,
- Interoception.M,
- ## Effectors
- Head.M,
- Mouth_throat.M,
- Torso.M,
- Hand_arm.M,
- Foot_leg.M
- )
- ),
- minkowski3_strength = sum(c(Vision.M, Hearing.M, Olfaction.M, Touch.M, Taste.M,
- Interoception.M, Head.M, Mouth_throat.M, Torso.M,
- Hand_arm.M, Foot_leg.M)^3)^(1/3),
- max_perceptual_strength = max(
- c(
- ## Modalities
- Vision.M,
- Hearing.M,
- Olfaction.M,
- Touch.M,
- Taste.M,
- Interoception.M
- )
- ),
- max_action_strength = max(
- c(
- ## Effectors
- Head.M,
- Mouth_throat.M,
- Torso.M,
- Hand_arm.M,
- Foot_leg.M
- )
- )
- ) %>%
- ungroup()
- df_dom_plus_sm %>%
- ggplot(aes(x = Max_strength.sensorimotor,
- y = max_strength)) +
- geom_point(alpha = .5) +
- labs(y = "Maximum Contextualized Strength",
- x = "Maximum Strength (Lancaster)") +
- theme_bw()
- cor.test(df_dom_plus_sm$Max_strength.sensorimotor,
- df_dom_plus_sm$max_strength)
- cor.test(df_dom_plus_sm$max_strength,
- df_dom_plus_sm$minkowski3_strength)
- ```
- ## Does contextualized sensorimotor strength predict dominance?
- The answer is **yes**: contexts with a higher *maximum* sensorimotor strength also tend to be rated as more *dominant*.
- Notably, this is true above and beyond the *decontextualized* ratings of sensorimotor strength for a given word.
- ```{r strength_dominance}
- df_dom_plus_sm %>%
- ggplot(aes(x = max_strength,
- y = dominance)) +
- geom_point(alpha = .4) +
- geom_smooth(method = "lm") +
- labs(x = "Maximum Contextualized Strength",
- y = "Dominance") +
- theme_minimal() +
- theme(text = element_text(size=16))
- df_dom_plus_sm %>%
- ggplot(aes(x = minkowski3_strength,
- y = dominance)) +
- geom_point(alpha = .4) +
- geom_smooth(method = "lm") +
- labs(x = "Minkowski-3 Sensorimotor Strength",
- y = "Dominance") +
- theme_minimal() +
- theme(text = element_text(size=16))
- model_full = lmer(data = df_dom_plus_sm,
- dominance ~
- max_strength +
- Max_strength.sensorimotor + Minkowski3.sensorimotor +
- (1 | word),
- REML = FALSE)
- model_reduced = lmer(data = df_dom_plus_sm,
- dominance ~
- # max_strength +
- Max_strength.sensorimotor + Minkowski3.sensorimotor +
- (1 | word),
- REML = FALSE)
- summary(model_full)
- anova(model_full, model_reduced)
- model_full = lmer(data = df_dom_plus_sm,
- dominance ~
- minkowski3_strength +
- Max_strength.sensorimotor + Minkowski3.sensorimotor +
- (1 | word),
- REML = FALSE)
- model_reduced = lmer(data = df_dom_plus_sm,
- dominance ~
- # max_strength +
- Max_strength.sensorimotor + Minkowski3.sensorimotor +
- (1 | word),
- REML = FALSE)
- summary(model_full)
- anova(model_full, model_reduced)
- ```
- ### Examples
- ```{r}
- ### Top 10
- df_dom_plus_sm %>%
- arrange(desc(minkowski3_strength)) %>%
- select(word, sentence, minkowski3_strength, dominance) %>%
- head(10)
- ### Bottom 10
- df_dom_plus_sm %>%
- arrange(minkowski3_strength) %>%
- select(word, sentence, minkowski3_strength, dominance) %>%
- head(10)
- ### Top 10
- df_dom_plus_sm %>%
- arrange(desc(max_strength)) %>%
- select(word, sentence, max_strength, dominance) %>%
- head(10)
- ### Bottom 10
- df_dom_plus_sm %>%
- arrange(max_strength) %>%
- select(word, sentence, max_strength, dominance) %>%
- head(10)
- ```
- ### Which dimension best predicts dominance?
- Here, we ask whether specific dimensions are particularly correlated with sense dominance.
- ```{r specific_dimensions_dominance}
- features <- c("Vision.M", "Hearing.M", "Olfaction.M", "Touch.M", "Taste.M",
- "Interoception.M", "Head.M", "Mouth_throat.M", "Torso.M",
- "Hand_arm.M", "Foot_leg.M")
- # Models with each feature + baseline covariate
- r2_results <- map_dfr(features, function(feat) {
- formula <- as.formula(paste0("dominance ~ Max_strength.sensorimotor + ", feat))
- model <- lm(formula, data = df_dom_plus_sm)
- tibble(
- feature = feat,
- R2 = summary(model)$r.squared,
- beta = coef(model)[3], # coefficient for the feature (3rd term now)
- p_value = summary(model)$coefficients[3, 4]
- )
- })
- # Baseline model (just Max_strength.sensorimotor)
- baseline_model <- lm(dominance ~ Max_strength.sensorimotor, data = df_dom_plus_sm)
- baseline_row <- tibble(
- feature = "Baseline",
- R2 = summary(baseline_model)$r.squared,
- beta = NA,
- p_value = NA
- )
- # Combine
- r2_results_with_baseline <- bind_rows(baseline_row, r2_results)
- # Plot
- r2_results_with_baseline %>%
- mutate(feature = str_remove(feature, "\\.M$"),
- feature = str_replace(feature, "_", "/"),
- sig = ifelse(p_value < .05, "*", ""),
- is_baseline = feature == "Baseline",
- feature = fct_reorder(feature, R2)) %>%
- ggplot(aes(x = R2, y = feature, fill = is_baseline)) +
- geom_col() +
- geom_text(aes(label = sig), hjust = -0.5, size = 6, na.rm = TRUE) +
- scale_fill_manual(values = c("grey40", "steelblue"), guide = "none") +
- labs(x = "R²", y = NULL, title = "Predicting Sense Dominance") +
- theme_minimal() +
- theme(text = element_text(size = 16))
- cor.test(df_dom_plus_sm$Touch.M,
- df_dom_plus_sm$dominance)
- cor.test(df_dom_plus_sm$Vision.M,
- df_dom_plus_sm$dominance)
- ```
- It looks like vision and touch are especially strong predictors. Which items drive this, e.g., for touch?
- ```{r touch_dominance}
- by_sense <- df_dom_plus_sm %>%
- mutate(sense = str_extract(context, "M[12]")) %>%
- group_by(word, sense) %>%
- summarize(
- mean_touch = mean(Touch.M),
- mean_vision = mean(Vision.M),
- mean_dom = mean(dominance),
- .groups = "drop"
- )
- # Now 2 rows per word—compute within-word difference
- sense_diffs <- by_sense %>%
- pivot_wider(names_from = sense,
- values_from = c(mean_touch, mean_vision, mean_dom)) %>%
- mutate(
- touch_diff = mean_touch_M1 - mean_touch_M2,
- vision_diff = mean_vision_M1 - mean_vision_M2,
- dom_diff = mean_dom_M1 - mean_dom_M2
- )
- sense_diffs %>%
- mutate(abs_touch_diff = abs(touch_diff)) %>%
- arrange(desc(abs_touch_diff)) %>%
- select(word, abs_touch_diff, dom_diff) %>%
- head(5)
- sense_diffs %>%
- mutate(abs_vision_diff = abs(vision_diff)) %>%
- arrange(desc(abs_vision_diff)) %>%
- select(word, abs_vision_diff, dom_diff) %>%
- head(5)
- # Does the sense with higher Touch or Vision tend to be more dominant?
- cor.test(sense_diffs$touch_diff, sense_diffs$dom_diff)
- cor.test(sense_diffs$vision_diff, sense_diffs$dom_diff)
- ```
- # How much does context add?
- Here, we inspect how much information the CS norms provide, both relative to the LS norms and just overall (i.e., how much context pushes sensorimotor vectors around).
- ## Comparing dimensions to Lancaster
- Now, we compare each dimension to the LS Norms.
- ```{r ls_deviation_market}
- df_diffs_raw <- df_dom_plus_sm %>%
- mutate(vision_diff = (Vision.M - Visual.mean),
- auditory_diff = (Hearing.M - Auditory.mean),
- intero_diff = (Interoception.M - Interoceptive.mean),
- olfactory_diff = (Olfaction.M - Olfactory.mean),
- touch_diff = (Touch.M - Haptic.mean),
- taste_diff = (Taste.M - Gustatory.mean),
- torso_diff = (Torso.M - Torso.mean),
- hand_arm_diff = (Hand_arm.M - Hand_arm.mean),
- foot_leg_diff = (Foot_leg.M - Foot_leg.mean),
- head_diff = (Head.M - Head.mean),
- mouth_throat_diff = (Mouth_throat.M - Mouth.mean)) %>%
- pivot_longer(cols = ends_with("_diff"),
- names_to = "Dimension", values_to = "Diff") %>%
- mutate(Dimension = gsub('_diff', '', Dimension),
- Dimension = case_when(
- Dimension == "intero" ~ "Interoception",
- Dimension == "auditory" ~ "Hearing",
- Dimension == "olfactory" ~ "Olfaction",
- TRUE ~ str_to_title(Dimension)
- ))
- df_diffs_raw$Dimension = factor(df_diffs_raw$Dimension,
- levels = rev(c(
- 'Vision',
- 'Hearing',
- 'Olfaction',
- 'Taste',
- 'Interoception',
- 'Touch',
- 'Mouth_throat',
- 'Head',
- 'Torso',
- 'Hand_arm',
- 'Foot_leg'
- )))
- df_diffs_raw %>%
- filter(word == "market") %>%
- ggplot(aes(x = Dimension,
- y = Diff,
- fill = Dimension)) +
- geom_bar(stat = "summary") +
- geom_vline(xintercept = 0, linetype = "dotted") +
- theme_bw() +
- coord_flip() +
- labs(x = "Dimension",
- y = "Deviation from Lancaster Norms") +
- scale_fill_manual(values = viridisLite::viridis(11, option = "mako",
- begin = 0.8, end = 0.15)) +
- facet_wrap(~sentence) +
- theme(text = element_text(size=16)) +
- guides(fill = FALSE)
- df_paired_market <- df_dom_plus_sm %>%
- filter(word == "market") %>%
- select(sentence,
- Vision.M, Hearing.M, Olfaction.M, Taste.M, Interoception.M, Touch.M,
- Mouth_throat.M, Head.M, Torso.M, Hand_arm.M, Foot_leg.M,
- Visual.mean, Auditory.mean, Olfactory.mean, Gustatory.mean,
- Interoceptive.mean, Haptic.mean, Mouth.mean, Head.mean,
- Torso.mean, Hand_arm.mean, Foot_leg.mean) %>%
- pivot_longer(cols = -sentence, names_to = "raw_dim", values_to = "Rating") %>%
- mutate(
- Source = case_when(
- str_detect(raw_dim, "\\.M$") ~ "CS Norms (contextualized)",
- TRUE ~ "LS Norms (decontextualized)"
- ),
- Dimension = case_when(
- raw_dim %in% c("Vision.M", "Visual.mean") ~ "Vision",
- raw_dim %in% c("Hearing.M", "Auditory.mean") ~ "Hearing",
- raw_dim %in% c("Olfaction.M", "Olfactory.mean") ~ "Olfaction",
- raw_dim %in% c("Taste.M", "Gustatory.mean") ~ "Taste",
- raw_dim %in% c("Interoception.M", "Interoceptive.mean") ~ "Interoception",
- raw_dim %in% c("Touch.M", "Haptic.mean") ~ "Touch",
- raw_dim %in% c("Mouth_throat.M", "Mouth.mean") ~ "Mouth_throat",
- raw_dim %in% c("Head.M", "Head.mean") ~ "Head",
- raw_dim %in% c("Torso.M", "Torso.mean") ~ "Torso",
- raw_dim %in% c("Hand_arm.M", "Hand_arm.mean") ~ "Hand_arm",
- raw_dim %in% c("Foot_leg.M", "Foot_leg.mean") ~ "Foot_leg"
- )
- )
- df_paired_market$Dimension <- factor(df_paired_market$Dimension,
- levels = rev(c('Vision','Hearing','Olfaction','Taste','Interoception',
- 'Touch','Mouth_throat','Head','Torso','Hand_arm','Foot_leg')))
- df_paired_market %>%
- ggplot(aes(x = Dimension, y = Rating, fill = Source)) +
- geom_bar(stat = "summary", position = position_dodge(width = 0.8), width = 0.7) +
- theme_bw() +
- coord_flip() +
- labs(x = "Dimension", y = "Sensorimotor Strength", fill = NULL) +
- scale_fill_manual(values = c("CS Norms (contextualized)" = "#2c7a7b",
- "LS Norms (decontextualized)" = "#7c5e9c")) +
- facet_wrap(~sentence) +
- theme(text = element_text(size = 14), legend.position = "bottom")
- ```
- ### Distribution of deviations
- ```{r largest_dev}
- df_diffs_raw %>%
- mutate(abs_diff = abs(Diff)) %>%
- summarise(mean_abs_diff = mean(abs_diff),
- median_abs_diff = median(abs_diff),
- sd_abs_diff = sd(abs_diff),
- min_abs_diff = min(abs_diff),
- max_abs_diff = max(abs_diff))
- df_diffs_raw %>%
- mutate(abs_diff = abs(Diff)) %>%
- ggplot(aes(x = abs_diff)) +
- geom_histogram() +
- labs(x = "Absolute Deviation") +
- theme_minimal() +
- facet_wrap(~Dimension) +
- theme(text = element_text(size=16))
- df_by_sentence_raw = df_diffs_raw %>%
- mutate(abs_diff = abs(Diff)) %>%
- group_by(word, sentence) %>%
- summarise(mean_abs_diff = mean(abs_diff),
- min_diff = min(abs_diff),
- max_diff = max(abs_diff)) %>%
- ungroup()
- df_by_sentence_raw %>%
- summarise(mean_diff = mean(mean_abs_diff),
- median_abs_diff = median(mean_abs_diff),
- sd_abs_diff = sd(mean_abs_diff),
- min_abs_diff = min(mean_abs_diff),
- max_abs_diff = max(mean_abs_diff))
- df_by_sentence_raw %>%
- ggplot(aes(x = mean_abs_diff)) +
- geom_histogram() +
- labs(x = "Mean Absolute Deviation (By Sentence)") +
- theme_minimal() +
- theme(text = element_text(size=16))
- df_by_sentence_raw %>%
- ggplot(aes(x = max_diff)) +
- geom_histogram() +
- labs(x = "Max Absolute Deviation (By Sentence)") +
- theme_minimal() +
- theme(text = element_text(size=16))
- df_by_sentence_raw %>%
- arrange(desc(max_diff)) %>%
- head(10)
- df_by_word_raw = df_diffs_raw %>%
- mutate(abs_diff = abs(Diff)) %>%
- group_by(word) %>%
- summarise(mean_abs_diff = mean(abs_diff)) %>%
- ungroup()
- df_by_word_raw %>%
- summarise(mean_diff = mean(mean_abs_diff),
- median_abs_diff = median(mean_abs_diff),
- sd_abs_diff = sd(mean_abs_diff),
- min_abs_diff = min(mean_abs_diff),
- max_abs_diff = max(mean_abs_diff))
- df_diffs_raw %>%
- mutate(abs_diff = abs(Diff)) %>%
- group_by(word) %>%
- summarise(mean_abs_diff = mean(abs_diff),
- sd_abs_diff = sd(abs_diff),
- range_diff = max(abs_diff) - min(abs_diff)) %>%
- arrange(desc(mean_abs_diff)) %>%
- head(5)
- ```
- ### Individual words
- ```{r individual_words}
- df_diffs_raw = df_diffs_raw %>%
- mutate(sense = str_extract(context, "M[12]"), # "M1" or "M2"
- context_within_sense = str_extract(context, "[ab]")) # "a" or "b"
- df_diffs_raw %>%
- filter(word == "jam") %>%
- mutate(sentence = fct_reorder(sentence, as.numeric(factor(sense)))) %>%
- ggplot(aes(x = Dimension,
- y = Diff,
- fill = Dimension)) +
- geom_bar(stat = "summary") +
- geom_vline(xintercept = 0, linetype = "dotted") +
- theme_bw() +
- coord_flip() +
- labs(x = "Dimension",
- y = "Deviation from Lancaster Norms") +
- scale_fill_manual(values = viridisLite::viridis(11, option = "mako",
- begin = 0.8, end = 0.15)) +
- facet_wrap(~sentence) +
- theme(text = element_text(size=16)) +
- guides(fill = FALSE)
- df_diffs_raw %>%
- mutate(sentence = fct_reorder(sentence, as.numeric(factor(sense)))) %>%
- filter(word == "punch") %>%
- ggplot(aes(x = Dimension,
- y = Diff,
- fill = Dimension)) +
- geom_bar(stat = "summary") +
- geom_vline(xintercept = 0, linetype = "dotted") +
- theme_bw() +
- coord_flip() +
- labs(x = "Dimension",
- y = "Deviation from Lancaster Norms") +
- scale_fill_manual(values = viridisLite::viridis(11, option = "mako",
- begin = 0.8, end = 0.15)) +
- facet_wrap(~sentence) +
- theme(text = element_text(size=16)) +
- guides(fill = FALSE)
- df_diffs_raw %>%
- mutate(sentence = fct_reorder(sentence, as.numeric(factor(sense)))) %>%
- filter(word == "match") %>%
- ggplot(aes(x = Dimension,
- y = Diff,
- fill = Dimension)) +
- geom_bar(stat = "summary") +
- geom_vline(xintercept = 0, linetype = "dotted") +
- theme_bw() +
- coord_flip() +
- labs(x = "Dimension",
- y = "Deviation from Lancaster Norms") +
- scale_fill_manual(values = viridisLite::viridis(11, option = "mako",
- begin = 0.8, end = 0.15)) +
- facet_wrap(~sentence) +
- theme(text = element_text(size=16)) +
- guides(fill = FALSE)
- df_diffs_raw %>%
- mutate(sentence = fct_reorder(sentence, as.numeric(factor(sense)))) %>%
- filter(word == "pen") %>%
- ggplot(aes(x = Dimension,
- y = Diff,
- fill = Dimension)) +
- geom_bar(stat = "summary") +
- geom_vline(xintercept = 0, linetype = "dotted") +
- theme_bw() +
- coord_flip() +
- labs(x = "Dimension",
- y = "Deviation from Lancaster Norms") +
- scale_fill_manual(values = viridisLite::viridis(11, option = "mako",
- begin = 0.8, end = 0.15)) +
- facet_wrap(~sentence) +
- theme(text = element_text(size=16)) +
- guides(fill = FALSE)
- df_diffs_raw %>%
- mutate(sentence = fct_reorder(sentence, as.numeric(factor(sense)))) %>%
- filter(word == "band") %>%
- ggplot(aes(x = Dimension,
- y = Diff,
- fill = Dimension)) +
- geom_bar(stat = "summary") +
- geom_vline(xintercept = 0, linetype = "dotted") +
- theme_bw() +
- coord_flip() +
- labs(x = "Dimension",
- y = "Deviation from Lancaster Norms") +
- scale_fill_manual(values = viridisLite::viridis(11, option = "mako",
- begin = 0.8, end = 0.15)) +
- facet_wrap(~sentence) +
- theme(text = element_text(size=16)) +
- guides(fill = FALSE)
- ```
- ### Which dimensions have the largest deviations?
- ```{r}
- df_diffs_raw %>%
- mutate(abs_diff = abs(Diff)) %>%
- group_by(Dimension) %>%
- summarise(mean_abs_diff = mean(abs_diff),
- sd_abs_diff = sd(abs_diff),
- range_diff = max(abs_diff) - min(abs_diff)) %>%
- arrange(desc(range_diff))
- ```
- ## Z-scored deviations
- ```{r zscored_diffs}
- SD_FLOOR <- 0.1
- df_diffs_z_floored <- df_dom_plus_sm %>%
- mutate(
- vision_diff = (Vision.M - Visual.mean) / pmax(Visual.SD, SD_FLOOR),
- auditory_diff = (Hearing.M - Auditory.mean) / pmax(Auditory.SD, SD_FLOOR),
- intero_diff = (Interoception.M - Interoceptive.mean) / pmax(Interoceptive.SD, SD_FLOOR),
- olfactory_diff = (Olfaction.M - Olfactory.mean) / pmax(Olfactory.SD, SD_FLOOR),
- touch_diff = (Touch.M - Haptic.mean) / pmax(Haptic.SD, SD_FLOOR),
- taste_diff = (Taste.M - Gustatory.mean) / pmax(Gustatory.SD, SD_FLOOR),
- torso_diff = (Torso.M - Torso.mean) / pmax(Torso.SD.LS, SD_FLOOR),
- hand_arm_diff = (Hand_arm.M - Hand_arm.mean) / pmax(Hand_arm.SD.LS, SD_FLOOR),
- foot_leg_diff = (Foot_leg.M - Foot_leg.mean) / pmax(Foot_leg.SD.LS, SD_FLOOR),
- head_diff = (Head.M - Head.mean) / pmax(Head.SD.LS, SD_FLOOR),
- mouth_throat_diff = (Mouth_throat.M - Mouth.mean) / pmax(Mouth.SD, SD_FLOOR)
- ) %>%
- pivot_longer(cols = ends_with("_diff"),
- names_to = "Dimension", values_to = "Diff") %>%
- mutate(Dimension = gsub('_diff', '', Dimension),
- Dimension = case_when(
- Dimension == "intero" ~ "Interoception",
- Dimension == "auditory" ~ "Hearing",
- Dimension == "olfactory" ~ "Olfaction",
- TRUE ~ str_to_title(Dimension)
- )) %>%
- select(word, sentence, Dimension, Diff, context, dominance)
- # Distribution
- sentence_dev_floored <- df_diffs_z_floored %>%
- mutate(abs_diff = abs(Diff)) %>%
- group_by(word, sentence) %>%
- summarise(mean_abs_diff = mean(abs_diff, na.rm = TRUE),
- max_abs_diff = max(abs_diff),
- .groups = "drop")
- sentence_dev_floored %>%
- ggplot(aes(x = mean_abs_diff)) +
- geom_histogram(bins = 30) +
- geom_vline(xintercept = 1, linetype = "dashed", color = "red") +
- labs(x = "Mean Absolute Standardized Deviation (By Sentence)",
- y = "Count of sentences") +
- theme_minimal() +
- theme(text = element_text(size=16))
- sentence_dev_floored %>%
- ggplot(aes(x = max_abs_diff)) +
- geom_histogram(bins = 30) +
- geom_vline(xintercept = 1, linetype = "dashed", color = "red") +
- labs(x = "Max Absolute Standardized Deviation (By Sentence)",
- y = "Count of sentences") +
- theme_minimal() +
- theme(text = element_text(size=16))
- ### Overall % sentences more than one on average
- sentence_dev_floored %>%
- mutate(greater_than_one = mean_abs_diff > 1) %>%
- summarise(prop_more_than_one = mean(greater_than_one),
- max_dev = max(mean_abs_diff))
- ### Overall % sentences more than one maximally
- quantile(sentence_dev_floored$max_abs_diff,
- probs = c(0.25, 0.5, 0.75, 0.9, 0.95, 0.99),
- na.rm = TRUE)
- ### Overall deviations by dimension
- df_diffs_z_floored %>%
- mutate(abs_diff = abs(Diff)) %>%
- ggplot(aes(x = abs_diff)) +
- geom_histogram() +
- labs(x = "Standardized Absolute Deviation") +
- geom_vline(xintercept = 1, linetype = "dashed", color = "red") +
- theme_minimal() +
- facet_wrap(~Dimension) +
- theme(text = element_text(size=16))
- ### Overall % of contexts exceeding 1 SD
- df_diffs_z_floored %>%
- mutate(abs_diff = abs(Diff)) %>%
- mutate(greater_than_one = abs_diff > 1) %>%
- summarise(prop_more_than_one = mean(greater_than_one),
- max_dev = max(abs_diff))
- #### Additional summary statistics
- df_diffs_z_floored %>%
- mutate(abs_diff = abs(Diff)) %>%
- summarise(mean_abs_diff = mean(abs_diff),
- median_abs_diff = median(abs_diff),
- sd_abs_diff = sd(abs_diff),
- min_abs_diff = min(abs_diff),
- max_abs_diff = max(abs_diff))
- df_by_sentence_z_floored = df_diffs_z_floored %>%
- mutate(abs_diff = abs(Diff)) %>%
- group_by(word, sentence) %>%
- summarise(mean_abs_diff = mean(abs_diff)) %>%
- ungroup()
- df_by_sentence_z_floored %>%
- summarise(mean_diff = mean(mean_abs_diff),
- median_abs_diff = median(mean_abs_diff),
- sd_abs_diff = sd(mean_abs_diff),
- min_abs_diff = min(mean_abs_diff),
- max_abs_diff = max(mean_abs_diff))
- df_by_word_z_floored = df_diffs_z_floored %>%
- mutate(abs_diff = abs(Diff)) %>%
- group_by(word) %>%
- summarise(mean_abs_diff = mean(abs_diff)) %>%
- ungroup()
- df_by_word_z_floored %>%
- summarise(mean_diff = mean(mean_abs_diff),
- median_abs_diff = median(mean_abs_diff),
- sd_abs_diff = sd(mean_abs_diff),
- min_abs_diff = min(mean_abs_diff),
- max_abs_diff = max(mean_abs_diff))
- ```
- ### Compare to raw differences
- ```{r}
- # Raw (unstandardized) mean absolute deviation per sentence
- sentence_dev_raw <- df_diffs_raw %>%
- mutate(abs_diff = abs(Diff)) %>%
- group_by(word, sentence, context) %>%
- summarise(mean_abs_diff = mean(abs_diff, na.rm = TRUE),
- .groups = "drop") %>%
- mutate(version = "Raw")
- # Standardized (floored) mean absolute deviation per sentence
- sentence_dev_z <- df_diffs_z_floored %>%
- mutate(abs_diff = abs(Diff)) %>%
- group_by(word, sentence, context) %>%
- summarise(mean_abs_diff = mean(abs_diff, na.rm = TRUE),
- .groups = "drop") %>%
- mutate(version = "Standardized")
- # By sentence
- sentence_compare <- sentence_dev_raw %>%
- select(word, sentence, context, raw = mean_abs_diff) %>%
- inner_join(sentence_dev_z %>%
- select(word, sentence, context, standardized = mean_abs_diff),
- by = c("word", "sentence", "context"))
- cor.test(sentence_compare$raw, sentence_compare$standardized, method = "spearman")
- cor.test(sentence_compare$raw, sentence_compare$standardized, method = "pearson")
- sentence_compare %>%
- ggplot(aes(x = raw, y = standardized)) +
- geom_point(alpha = 0.3) +
- geom_smooth(method = "lm", se = FALSE) +
- labs(x = "Raw mean absolute deviation",
- y = "Standardized mean absolute deviation",
- title = "By-sentence deviations") +
- theme_minimal() +
- theme(text = element_text(size=16))
- ```
- ## Within-word range
- Here, we calculate the *range* of sensorimotor associatoins for each dimension across the contexts of use for each word.
- ### Across all contexts
- ```{r within_word_range}
- df_rawc_with_norms = read_csv("../../data/processed/sentence_pairs_with_sensorimotor_distance.csv") %>%
- drop_na(sensorimotor_distance) %>%
- select(sensorimotor_distance, action_distance, perceptual_distance,
- word, same, ambiguity_type, sentence1, sentence2, mean_relatedness, Class, version)
- nrow(df_rawc_with_norms)
- word_dim_range <- df_dom_plus_sm %>%
- pivot_longer(cols = c(Vision.M, Hearing.M, Olfaction.M, Taste.M,
- Interoception.M, Touch.M, Mouth_throat.M,
- Head.M, Hand_arm.M, Foot_leg.M, Torso.M),
- names_to = "Dimension", values_to = "CS_value") %>%
- mutate(Dimension = gsub("\\.M$", "", Dimension)) %>%
- group_by(word, Dimension) %>%
- summarise(within_range = max(CS_value, na.rm = TRUE) - min(CS_value, na.rm = TRUE),
- .groups = "drop")
- # Aggregate to word level (mean range across dimensions)
- word_range <- word_dim_range %>%
- group_by(word) %>%
- summarise(mean_range = mean(within_range, na.rm = TRUE),
- max_range = max(within_range, na.rm = TRUE),
- .groups = "drop") %>%
- arrange(desc(mean_range))
- print(head(word_range, 15))
- print(tail(word_range, 15))
- # By dimension
- word_dim_range %>%
- ggplot(aes(x = within_range)) +
- geom_histogram(bins = 30) +
- labs(x = "Within-Word Range of CS Ratings (across contexts)",
- y = "Count of words") +
- theme_minimal() +
- facet_wrap(~Dimension)
- # Distribution
- word_range %>%
- ggplot(aes(x = mean_range)) +
- geom_histogram(bins = 30) +
- labs(x = "Mean Within-Word Range of CS Ratings (across contexts)",
- y = "Count of words") +
- theme_minimal()
- # Summary stats
- word_range %>%
- summarise(mean = mean(mean_range),
- sd = sd(mean_range),
- median = median(mean_range),
- min = min(mean_range),
- max = max(mean_range))
- # Get per-word ambiguity type
- word_ambiguity <- df_rawc_with_norms %>%
- distinct(word, ambiguity_type)
- # Merge with within-word range
- word_range_with_amb <- word_range %>%
- inner_join(word_ambiguity, by = "word")
- # Summary stats by ambiguity type
- word_range_with_amb %>%
- group_by(ambiguity_type) %>%
- summarise(mean = mean(mean_range),
- sd = sd(mean_range),
- median = median(mean_range),
- n = n(),
- .groups = "drop")
- # Test the difference
- summary(lm(data = word_range_with_amb,
- mean_range ~ ambiguity_type))
- t.test(mean_range ~ ambiguity_type, data = word_range_with_amb)
- # Visualize
- word_range_with_amb %>%
- ggplot(aes(x = ambiguity_type, y = mean_range, fill = ambiguity_type)) +
- geom_boxplot(alpha = 0.6) +
- geom_jitter(width = 0.15, alpha = 0.3) +
- labs(x = "Ambiguity Type",
- y = "Mean Within-Word Range of CS Ratings") +
- theme_minimal() +
- theme(legend.position = "none")
- ```
- ### Between vs. Within sense
- ```{r within_word_range_senses}
- # Need sense info from context column (M1/M2)
- word_range_decomposed <- df_dom_plus_sm %>%
- mutate(sense = str_extract(context, "M[12]")) %>%
- pivot_longer(cols = c(Vision.M, Hearing.M, Olfaction.M, Taste.M,
- Interoception.M, Touch.M, Mouth_throat.M,
- Head.M, Hand_arm.M, Foot_leg.M, Torso.M),
- names_to = "Dimension", values_to = "CS_value") %>%
- mutate(Dimension = gsub("\\.M$", "", Dimension)) %>%
- group_by(word, Dimension, sense) %>%
- summarise(sense_mean = mean(CS_value),
- sense_within_range = max(CS_value) - min(CS_value),
- .groups = "drop") %>%
- group_by(word, Dimension) %>%
- summarise(between_sense_range = max(sense_mean) - min(sense_mean),
- within_sense_range = mean(sense_within_range),
- .groups = "drop") %>%
- group_by(word) %>%
- summarise(mean_between_sense = mean(between_sense_range),
- mean_within_sense = mean(within_sense_range),
- max_between_sense = max(between_sense_range),
- max_within_sense = max(within_sense_range),
- .groups = "drop")
- word_range_decomposed %>%
- inner_join(word_ambiguity, by = "word") %>%
- group_by(ambiguity_type) %>%
- summarise(between = mean(mean_between_sense),
- within = mean(mean_within_sense),,
- between_max = mean(max_between_sense),
- within_max = mean(max_within_sense),
- n = n())
- # Test each separately
- t.test(mean_between_sense ~ ambiguity_type,
- data = word_range_decomposed %>% inner_join(word_ambiguity, by = "word"))
- t.test(mean_within_sense ~ ambiguity_type,
- data = word_range_decomposed %>% inner_join(word_ambiguity, by = "word"))
- # Test max
- t.test(max_between_sense ~ ambiguity_type,
- data = word_range_decomposed %>% inner_join(word_ambiguity, by = "word"))
- t.test(max_within_sense ~ ambiguity_type,
- data = word_range_decomposed %>% inner_join(word_ambiguity, by = "word"))
- word_range_decomposed_long = word_range_decomposed %>%
- inner_join(word_ambiguity, by = "word") %>%
- pivot_longer(cols = c(mean_between_sense, mean_within_sense),
- names_to = "comparison",
- values_to = "word_range")
- summary(lmer(data = word_range_decomposed_long,
- word_range ~ comparison + (1|word)))
- summary(lmer(data = word_range_decomposed_long,
- word_range ~ ambiguity_type*comparison + (1|word)))
- word_range_decomposed_long %>%
- mutate(comparison = case_when(
- comparison == "mean_between_sense" ~ "Different Sense",
- comparison == "mean_within_sense" ~ "Same Sense"
- )) %>%
- ggplot(aes(x = word_range,
- y = ambiguity_type,
- fill = comparison)) +
- geom_density_ridges2(aes(height = ..density..),
- color = NA,
- alpha = 0.5,
- scale=0.85,
- stat="density") +
- labs(x = "Mean Within-Word Range of CS Ratings",
- y = NULL,
- fill = NULL) +
- theme_minimal() +
- scale_fill_manual(
- values = c("Same Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[1],
- "Different Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[2])
- ) +
- theme(text = element_text(size = 20))
- word_range_decomposed_long = word_range_decomposed %>%
- inner_join(word_ambiguity, by = "word") %>%
- pivot_longer(cols = c(max_between_sense, max_within_sense),
- names_to = "comparison",
- values_to = "word_range")
- summary(lmer(data = word_range_decomposed_long,
- word_range ~ comparison + (1|word)))
- summary(lmer(data = word_range_decomposed_long,
- word_range ~ ambiguity_type*comparison + (1|word)))
- word_range_decomposed_long %>%
- mutate(comparison = case_when(
- comparison == "max_between_sense" ~ "Different Sense",
- comparison == "max_within_sense" ~ "Same Sense"
- )) %>%
- ggplot(aes(x = word_range,
- y = ambiguity_type,
- fill = comparison)) +
- geom_density_ridges2(aes(height = ..density..),
- color = NA,
- alpha = 0.5,
- scale=0.85,
- stat="density") +
- labs(x = "Max Within-Word Range of CS Ratings",
- y = NULL,
- fill = NULL) +
- theme_minimal() +
- scale_fill_manual(
- values = c("Same Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[1],
- "Different Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[2])
- ) +
- theme(text = element_text(size = 20))
- ```
- # How does dominance relate to deviation from the LS Norms?
- This question can in turn be decomposed into two questions:
- First, are more dominant senses **closer** to the LS Norms overall? We might expect this to be the case if the LS Norms reflect the dominant sense; that is, when people rate the sensorimotor properties of a decontextualized word, they might be more likely to index properties associated with the most dominant contexts or meanings of that word.
- And the answer is **yes**: more dominant senses are indeed more *similar* (less distant) from the Lancaster norm in terms of their sensorimotor profile.
- ```{r dominance_ls_deviation}
- model_with_dominance = lmer(data = df_dom_plus_sm,
- distance_to_lancaster ~ dominance + (1 | word),
- REML = FALSE)
- model_no_dominance = lmer(data = df_dom_plus_sm,
- distance_to_lancaster ~ (1 | word),
- REML = FALSE)
- summary(model_with_dominance)
- anova(model_with_dominance, model_no_dominance)
- df_dom_plus_sm %>%
- ggplot(aes(x = dominance,
- y = distance_to_lancaster)) +
- geom_point(alpha = .5) +
- geom_smooth(method = "lm") +
- labs(x = "Dominance",
- y = "Cosine Distance to Decontextualized LS Norm") +
- theme_bw()
- cor.test(df_dom_plus_sm$dominance, df_dom_plus_sm$distance_to_lancaster)
- ```
- And second: does dominance predict the **direction** of difference?
- The earlier analysis of dominance suggests that more dominant senses are more concrete than less dominant senses. Thus, we might expect that more dominant senses are also more concrete on average than the decontextualized norms.
- We find that this is true: that is, more dominant senses are *more concrete* on average than the LS norm.
- ```{r}
- df_diffs_avg = df_diffs_raw %>%
- group_by(word, sentence, context) %>%
- summarise(mean_diff = mean(Diff))
- df_diffs_avg = df_diffs_avg %>%
- left_join(df_dom_plus_sm)
- model_with_dominance = lmer(data = df_diffs_avg,
- mean_diff ~ dominance + (1 | word),
- REML = FALSE)
- model_no_dominance = lmer(data = df_diffs_avg,
- mean_diff ~ (1 | word),
- REML = FALSE)
- summary(model_with_dominance)
- anova(model_with_dominance, model_no_dominance)
- cor.test(df_diffs_avg$dominance, df_diffs_avg$mean_diff)
- ```
- # Predicting relatedness
- Next, we ask about the **sensorimotor distance** between two sentence pairs, and whether it correlates both with `same/different sense` and the `mean_relatedness` judgments for those sentence pairs.
- Here, we load a version of the dataset that also contains a *baseline* measure: the sensorimotor distance as calculated using a bag-of-words approach (i.e., using the original Lancaster Norms).
- ## Load data
- ```{r}
- df_rawc_with_norms = read_csv("../../data/processed/sentence_pairs_with_sensorimotor_distance.csv") %>%
- drop_na(sensorimotor_distance) %>%
- select(sensorimotor_distance, action_distance, perceptual_distance,
- word, same, ambiguity_type, sentence1, sentence2, mean_relatedness, Class, version)
- nrow(df_rawc_with_norms)
- ```
- ## Load English RAW-C data
- ```{r}
- df_bert = read_csv("../../data/processed/models_english/rawc-distances_model-bert-base-uncased.csv") %>%
- mutate(Model = "BERT-base-uncased",
- Multilingual = "Monolingual")
- df_bert_cased = read_csv("../../data/processed/models_english/rawc-distances_model-bert-base-cased.csv") %>%
- mutate(Model = "BERT-base-cased",
- Multilingual = "Monolingual")
- df_xlm = read_csv("../../data/processed/models_english/rawc-distances_model-xlm-roberta-base.csv") %>%
- mutate(Model = "XLM-RoBERTa",
- Multilingual = "Multilingual")
- df_ab1 = read_csv("../../data/processed/models_english/rawc-distances_model-albert-base-v1.csv") %>%
- mutate(Model = "ALBERT-base-v1",
- Multilingual = "Monolingual")
- df_ab2 = read_csv("../../data/processed/models_english/rawc-distances_model-albert-base-v2.csv") %>%
- mutate(Model = "ALBERT-base-v2",
- Multilingual = "Monolingual")
- df_al = read_csv("../../data/processed/models_english/rawc-distances_model-albert-large-v2.csv") %>%
- mutate(Model = "ALBERT-large-v2",
- Multilingual = "Monolingual")
- df_axl = read_csv("../../data/processed/models_english/rawc-distances_model-albert-xlarge-v2.csv") %>%
- mutate(Model = "ALBERT-xlarge-v2",
- Multilingual = "Monolingual")
- df_axxl = read_csv("../../data/processed/models_english/rawc-distances_model-albert-xxlarge-v2.csv") %>%
- mutate(Model = "ALBERT-xxlarge-v2",
- Multilingual = "Monolingual")
- df_rb = read_csv("../../data/processed/models_english/rawc-distances_model-roberta-base.csv") %>%
- mutate(Model = "RoBERTa-base",
- Multilingual = "Monolingual")
- df_rl = read_csv("../../data/processed/models_english/rawc-distances_model-roberta-large.csv") %>%
- mutate(Model = "RoBERTa-large",
- Multilingual = "Monolingual")
- df_db = read_csv("../../data/processed/models_english/rawc-distances_model-distilbert-base-uncased.csv") %>%
- mutate(Model = "DistilBERT",
- Multilingual = "Monolingual")
- df_mb = read_csv("../../data/processed/models_english/rawc-distances_model-bert-base-multilingual-cased.csv") %>%
- mutate(Model = "Multilingual BERT",
- Multilingual = "Multilingual")
- df_all = df_bert %>%
- bind_rows(df_bert_cased) %>%
- bind_rows(df_xlm) %>%
- bind_rows(df_ab1) %>%
- bind_rows(df_ab2) %>%
- bind_rows(df_al) %>%
- bind_rows(df_axl) %>%
- bind_rows(df_axxl) %>%
- bind_rows(df_rb) %>%
- bind_rows(df_rl) %>%
- bind_rows(df_db) %>%
- bind_rows(df_mb)
- df_merged = df_rawc_with_norms %>%
- inner_join(df_all)
- ```
- ## Predicting same/different sense
- ### Sensorimotor distance and sense boundaries
- ```{r sm_sense}
- df_rawc_with_norms = df_rawc_with_norms %>%
- mutate(Same = case_when(
- same == TRUE ~ "Same Sense",
- same == FALSE ~ "Different Sense"
- ))
- df_rawc_with_norms %>%
- ggplot(aes(x = sensorimotor_distance,
- y = ambiguity_type,
- fill = Same)) +
- geom_density_ridges2(aes(height = ..density..),
- color = NA,
- alpha = 0.5,
- scale=0.85,
- stat="density") +
- labs(x = "Sensorimotor Distance",
- y = NULL,
- fill = NULL) +
- scale_fill_manual(
- values = c("Same Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[1],
- "Different Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[2])
- ) +
- theme_minimal() +
- theme(text = element_text(size = 20))
- model_full = lmer(data = df_rawc_with_norms,
- sensorimotor_distance ~ same +
- (1 + same | word),
- control=lmerControl(optimizer="bobyqa"),
- REML = FALSE)
- model_reduced = lmer(data = df_rawc_with_norms,
- sensorimotor_distance ~ # same +
- (1 + same | word),
- control=lmerControl(optimizer="bobyqa"),
- REML = FALSE)
- summary(model_full)
- anova(model_full, model_reduced)
- model_full_at = lmer(data = df_rawc_with_norms,
- sensorimotor_distance ~ same * ambiguity_type +
- (1 + same | word),
- control=lmerControl(optimizer="bobyqa"),
- REML = FALSE)
- model_reduced_at = lmer(data = df_rawc_with_norms,
- sensorimotor_distance ~ same + ambiguity_type +
- (1 + same | word),
- control=lmerControl(optimizer="bobyqa"),
- REML = FALSE)
- summary(model_full_at)
- anova(model_full_at, model_reduced_at)
- df_rawc_with_norms %>%
- group_by(same, ambiguity_type) %>%
- summarise(mean = mean(sensorimotor_distance),
- sd = sd(sensorimotor_distance))
- ```
- ### Predicting sense boundaries
- Here, we run the actual analysis:
- ```{r aic_sense_boundaries}
- model_sm_only <- glm(same ~ sensorimotor_distance,
- data = df_rawc_with_norms,
- family = binomial)
- aic_sm_baseline <- AIC(model_sm_only)
- print(paste("AIC:", round(aic_sm_baseline, 1)))
- ### Get aic
- aic_results <- df_merged %>%
- group_by(n_params, Model, Layer) %>%
- summarize(
- # BERT only
- aic_bert = {
- model <- glm(Same_sense ~ Distance, family = binomial)
- AIC(model)
- },
- # Combined
- aic_combined = {
- model <- glm(Same_sense ~ Distance + sensorimotor_distance, family = binomial)
- AIC(model)
- },
- aic_sm = {
- model <- glm(Same_sense ~ sensorimotor_distance, family = binomial)
- AIC(model)
- },
- .groups = "drop"
- ) %>%
- mutate(
- # AIC Delta
- aic_bert_vs_sm = aic_bert - aic_sm_baseline,
- aic_combined_vs_sm = aic_sm_baseline - aic_combined,
- aic_combined_vs_dist = aic_bert - aic_combined
- )
- # ============================================================
- # 3. Best layer per model
- # ============================================================
- best_layers <- aic_results %>%
- group_by(Model, n_params) %>%
- slice_min(aic_bert, n = 1) %>%
- select(Model, Layer,
- aic_bert_vs_sm, aic_combined_vs_sm, aic_combined_vs_dist,
- aic_sm, aic_bert, aic_combined)
- best_layers = best_layers %>%
- mutate(aic_bert_vs_sm2 = aic_bert - aic_sm_baseline)
- print(best_layers)
- ## AIC difference from SM baseline
- best_layers %>%
- select(Model, aic_bert_vs_sm, aic_combined_vs_sm) %>%
- pivot_longer(cols = c(aic_bert_vs_sm, aic_combined_vs_sm),
- names_to = "type", values_to = "aic_diff") %>%
- mutate(type = ifelse(type == "aic_bert_vs_sm",
- "Distributional only",
- "Distributional + Sensorimotor"),
- type = factor(type, levels = c("Distributional only",
- "Distributional + Sensorimotor")),
- Model = fct_reorder(Model, aic_diff)) %>%
- ggplot(aes(x = aic_diff, y = Model, fill = type)) +
- geom_col(position = "dodge") +
- geom_vline(xintercept = 0, linetype = "dashed", color = "red") +
- labs(x = "ΔAIC vs. Sensorimotor baseline",
- y = NULL,
- fill = NULL) +
- scale_fill_manual(values = c("Distributional only" = "gray60",
- "Distributional + Sensorimotor" = "steelblue")) +
- theme_minimal() +
- theme(text = element_text(size = 16),
- legend.position = "bottom")
- best_layers %>%
- ungroup() %>%
- select(Model, aic_bert, aic_combined) %>%
- mutate(across(starts_with("aic"), round)) %>%
- kbl(col.names = c("Model", "Distributional", "Hybrid"),
- format = "latex",
- booktabs = TRUE,
- caption = "AIC values for predicting sense boundaries using Distributional Distance alone or both Distributional Distance and Sensorimotor Distance (AIC for Sensorimotor Distance was 707)",
- label = "aic") %>%
- kable_styling(latex_options = "hold_position")
- aic_results %>%
- mutate(dist_better_than_sm = aic_bert_vs_sm > 4,
- hybrid_better_than_sm = aic_combined_vs_sm > 4,
- hybrid_better_than_dist = aic_combined_vs_dist > 4) %>%
- summarise(mean(dist_better_than_sm),
- mean(hybrid_better_than_sm),
- mean(hybrid_better_than_dist))
- best_layers %>%
- ungroup() %>%
- mutate(dist_better_than_sm = aic_bert_vs_sm < -4,
- hybrid_better_than_sm = aic_combined_vs_sm > 4,
- hybrid_better_than_dist = aic_combined_vs_dist > 4) %>%
- summarise(mean(dist_better_than_sm),
- mean(hybrid_better_than_sm),
- mean(hybrid_better_than_dist))
- ### Visualize raw AIC
- aic_summary_full <- best_layers %>%
- select(Model, Layer, aic_sm, aic_bert, aic_combined) %>%
- pivot_longer(cols = c(aic_sm, aic_bert, aic_combined),
- names_to = "type", values_to = "aic") %>%
- mutate(type = case_when(
- type == "aic_sm" ~ "Sensorimotor",
- type == "aic_bert" ~ "Distributional",
- type == "aic_combined" ~ "Hybrid"
- ),
- type = factor(type, levels = c("Sensorimotor", "Distributional", "Hybrid")))
- # Get sensorimotor baseline (should be the same for all models)
- aic_sm_baseline <- aic_summary_full %>%
- filter(type == "Sensorimotor") %>%
- pull(aic) %>%
- unique()
- # Filter to just distributional and hybrid
- aic_summary_no_sm <- aic_summary_full %>%
- filter(type != "Sensorimotor")
- aic_summary_se_no_sm <- aic_summary_no_sm %>%
- group_by(type) %>%
- summarize(
- mean_aic = mean(aic),
- se_aic = sd(aic) / sqrt(n())
- )
- ggplot() +
- # Sensorimotor baseline as dashed line
- geom_hline(yintercept = aic_sm_baseline,
- linetype = "dashed", color = "coral", linewidth = 1) +
- # Raw points for each model/layer
- geom_jitter(data = aic_summary_no_sm,
- aes(x = type, y = aic, color = type),
- alpha = 0.3, width = 0.1, size = 2) +
- # Mean + SE
- geom_point(data = aic_summary_se_no_sm,
- aes(x = type, y = mean_aic, color = type),
- size = 5) +
- geom_errorbar(data = aic_summary_se_no_sm,
- aes(x = type, ymin = mean_aic - se_aic,
- ymax = mean_aic + se_aic, color = type),
- width = 0.15, linewidth = 1.2) +
- annotate("text", x = 1.5, y = aic_sm_baseline + 30,
- label = "Sensorimotor", color = "coral", size = 5)+
- labs(x = NULL,
- y = "AIC (lower is better)",
- title = "Predicting Sense Boundary") +
- scale_color_manual(values = c("Distributional" = "gray60",
- "Hybrid" = "steelblue")) +
- theme_minimal() +
- theme(text = element_text(size = 16),
- legend.position = "none")
- ```
- ## Predicting relatedness
- ### Comparing to each distributional relatedness measure
- ```{r relatedness_comparison}
- df_rawc_with_norms %>%
- ggplot(aes(x = sensorimotor_distance,
- y = mean_relatedness,
- color = Same)) +
- geom_point(alpha = .5) +
- scale_color_manual(
- values = c("Same Sense" = viridisLite::viridis(2, option = "mako",
- begin = 0.8, end = 0.15)[1],
- "Different Sense" = viridisLite::viridis(2, option = "mako",
- begin = 0.8, end = 0.15)[2])
- ) +
- theme_minimal() +
- labs(x = "Sensorimotor Distance",
- y = "Mean Relatedness",
- color = "") +
- theme(text = element_text(size = 15),
- legend.position="bottom")
- cor(df_rawc_with_norms$sensorimotor_distance,
- df_rawc_with_norms$mean_relatedness)
- ```
- ### Comparing AIC
- ```{r aic_relatedness}
- # ============================================================
- # 2. For each model/layer: compute R² and AIC, difference from SM baseline
- # ============================================================
- relatedness_results <- df_merged %>%
- group_by(Model, Layer, n_params, Multilingual) %>%
- summarize(
- # Sensorimotor only
- r2_sm = {
- model <- lm(mean_relatedness ~ sensorimotor_distance)
- summary(model)$r.squared
- },
- aic_sm = {
- model <- lm(mean_relatedness ~ sensorimotor_distance)
- AIC(model)
- },
- # BERT only
- r2_bert = {
- model <- lm(mean_relatedness ~ Distance)
- summary(model)$r.squared
- },
- aic_bert = {
- model <- lm(mean_relatedness ~ Distance)
- AIC(model)
- },
- # Combined
- r2_combined = {
- model <- lm(mean_relatedness ~ Distance + sensorimotor_distance)
- summary(model)$r.squared
- },
- aic_combined = {
- model <- lm(mean_relatedness ~ Distance + sensorimotor_distance)
- AIC(model)
- },
- .groups = "drop"
- ) %>%
- mutate(
- aic_bert_vs_sm = aic_bert - aic_sm,
- aic_combined_vs_sm = aic_sm - aic_combined,
- aic_combined_vs_dist = aic_bert - aic_combined
- )
- # ============================================================
- # 3. Best layer per model
- # ============================================================
- best_layers <- relatedness_results %>%
- group_by(Model, n_params, Multilingual) %>%
- slice_max(r2_bert, n = 1) %>%
- select(Model, Layer, n_params, Multilingual,
- # r2_sm, r2_bert, r2_combined,
- aic_bert_vs_sm, aic_sm, aic_bert, aic_combined,
- aic_combined_vs_sm, aic_combined_vs_dist)
- print(best_layers)
- best_layers %>%
- ungroup() %>%
- select(Model, aic_bert, aic_combined) %>%
- mutate(across(starts_with("aic"), round)) %>%
- kbl(col.names = c("Model", "Distributional", "Hybrid"),
- format = "latex",
- booktabs = TRUE,
- caption = "AIC values for predicting relatedness judgments using Distributional Distance alone or both Distributional Distance and Sensorimotor Distance (AIC for Sensorimotor Distance was 2134.78.)",
- label = "aic") %>%
- kable_styling(latex_options = "hold_position")
- best_layers %>%
- select(Model, aic_bert_vs_sm, aic_combined_vs_sm) %>%
- pivot_longer(cols = c(aic_bert_vs_sm, aic_combined_vs_sm),
- names_to = "type", values_to = "aic_diff") %>%
- mutate(type = ifelse(type == "aic_bert_vs_sm", "Distributional only", "Distributional + Sensorimotor"),
- type = factor(type, levels = c("Distributional only", "Distributional + Sensorimotor")),
- Model = fct_reorder(Model, aic_diff)) %>%
- ggplot(aes(x = aic_diff, y = Model, fill = type)) +
- geom_col(position = "dodge") +
- geom_vline(xintercept = 0, linetype = "dashed", color = "red") +
- geom_vline(xintercept = 4, linetype = "dotted", color = "gray50") +
- geom_vline(xintercept = -4, linetype = "dotted", color = "gray50") +
- labs(x = "ΔAIC vs. Sensorimotor baseline",
- y = NULL,
- fill = NULL) +
- scale_fill_manual(values = c("Distributional only" = "gray60",
- "Distributional + Sensorimotor" = "steelblue")) +
- theme_minimal() +
- theme(text = element_text(size = 16),
- legend.position = "bottom")
- # ============================================================
- # 7. Summary
- # ============================================================
- best_layers %>%
- ungroup() %>%
- mutate(dist_better_than_sm = aic_bert_vs_sm < -4,
- hybrid_better_than_sm = aic_combined_vs_sm > 4,
- hybrid_better_than_dist = aic_combined_vs_dist > 4) %>%
- summarise(mean(dist_better_than_sm),
- mean(hybrid_better_than_sm),
- mean(hybrid_better_than_dist))
- ### Visualize raw AIC
- aic_summary_full <- best_layers %>%
- select(Model, Layer, aic_sm, aic_bert, aic_combined) %>%
- pivot_longer(cols = c(aic_sm, aic_bert, aic_combined),
- names_to = "type", values_to = "aic") %>%
- mutate(type = case_when(
- type == "aic_sm" ~ "Sensorimotor",
- type == "aic_bert" ~ "Distributional",
- type == "aic_combined" ~ "Hybrid"
- ),
- type = factor(type, levels = c("Sensorimotor", "Distributional", "Hybrid")))
- # Get sensorimotor baseline (should be the same for all models)
- aic_sm_baseline <- aic_summary_full %>%
- filter(type == "Sensorimotor") %>%
- pull(aic) %>%
- unique()
- # Filter to just distributional and hybrid
- aic_summary_no_sm <- aic_summary_full %>%
- filter(type != "Sensorimotor")
- aic_summary_se_no_sm <- aic_summary_no_sm %>%
- group_by(type) %>%
- summarize(
- mean_aic = mean(aic),
- se_aic = sd(aic) / sqrt(n())
- )
- ggplot() +
- # Sensorimotor baseline as dashed line
- geom_hline(yintercept = aic_sm_baseline,
- linetype = "dashed", color = "coral", linewidth = 1) +
- # Raw points for each model/layer
- geom_jitter(data = aic_summary_no_sm,
- aes(x = type, y = aic, color = type),
- alpha = 0.3, width = 0.1, size = 2) +
- # Mean + SE
- geom_point(data = aic_summary_se_no_sm,
- aes(x = type, y = mean_aic, color = type),
- size = 5) +
- geom_errorbar(data = aic_summary_se_no_sm,
- aes(x = type, ymin = mean_aic - se_aic,
- ymax = mean_aic + se_aic, color = type),
- width = 0.15, linewidth = 1.2) +
- annotate("text", x = 1.5, y = aic_sm_baseline + 30,
- label = "Sensorimotor", color = "coral", size = 5)+
- labs(x = NULL,
- y = "AIC (lower is better)",
- title = "Predicting Relatedness") +
- scale_color_manual(values = c("Distributional" = "gray60",
- "Hybrid" = "steelblue")) +
- theme_minimal() +
- theme(text = element_text(size = 16),
- legend.position = "none")
- ```
- ## Correlation between sensorimotor and distributional distance
- ```{r corr_dist_sm}
- # Get best layer
- df_by_layer = df_merged %>%
- group_by(Model, Multilingual, Layer, n_params) %>%
- summarise(r = cor(mean_relatedness, Distance, method = "pearson"),
- r2 = r ** 2,
- rho = cor(mean_relatedness, Distance, method = "spearman"),
- count = n())
- df_best_layer <- df_by_layer %>%
- group_by(Model) %>%
- slice_max(r2, n = 1) %>%
- select(Model, Layer, r2)
- # Filter for only best layers
- df_best <- df_merged %>%
- semi_join(df_best_layer, by = c("Model", "Layer"))
- # Pivot wider
- df_wide <- df_best %>%
- mutate(`Sensorimotor Distance` = sensorimotor_distance) %>%
- select(word, sentence1, sentence2, `Sensorimotor Distance`, Model, Distance) %>%
- pivot_wider(names_from = Model, values_from = Distance)
- # Compute correlation matrix
- cor_matrix <- df_wide %>%
- select(`Sensorimotor Distance`, where(is.numeric)) %>%
- cor(use = "pairwise.complete.obs")
- # Plot the correlation matrix
- ggcorrplot(cor_matrix,
- hc.order = FALSE,
- method = "square") +
- theme(
- axis.text.x = element_text(size = 10, angle = 45, hjust = 1),
- axis.text.y = element_text(size = 10)
- )
- ```
- # Baseline with LS Norms
- ```{r}
- df_with_baseline = read_csv("../../data/processed/sentence_pairs_with_baseline.csv") %>%
- drop_na(sensorimotor_distance) %>%
- drop_na(baseline_distance)
- nrow(df_with_baseline)
- ```
- ## Does sensorimotor distance predict relatedness above the baseline?
- We also ask whether whether our measure of contextualized sensorimotor distance predicts relatedness above and beyond a baseline that simply considers the decontextualized Lancaster Sensorimotor Norms for the dismabiguating words in a sentence. (We find that it does.)
- ```{r}
- model_bow_sm = lmer(data = df_with_baseline,
- mean_relatedness ~ baseline_distance + sensorimotor_distance +
- (1| word),
- REML = FALSE)
- model_just_bow = lmer(data = df_with_baseline,
- mean_relatedness ~ baseline_distance +
- (1| word),
- REML = FALSE)
- model_just_sm = lmer(data = df_with_baseline,
- mean_relatedness ~ sensorimotor_distance +
- (1| word),
- REML = FALSE)
- summary(model_bow_sm)
- anova(model_bow_sm, model_just_bow)
- anova(model_bow_sm, model_just_sm)
- ```
- # Predicting Trott & Bergen (2023)
- Here, we also ask whether sensorimotor distance can account for variance in RT and accuracy on a primed sensibility judgment task.
- ```{r}
- # ============================================================
- # Load T&B 2023 data and merge with norms
- # ============================================================
- df_s1 <- read_csv("../../data/tb2023/polysemy_s1_final.csv")
- df_s2 <- read_csv("../../data/tb2023/polysemy_s2_final.csv")
- df_tb2023 <- bind_rows(df_s1, df_s2)
- nrow(df_tb2023)
- # Item-level aggregates
- df_tb2023_agg_acc <- df_tb2023 %>%
- group_by(word, same, version) %>%
- summarise(accuracy = mean(correct_response), .groups = "drop")
- df_tb2023_agg_rt <- df_tb2023 %>%
- filter(correct_response == TRUE) %>%
- group_by(word, same, version) %>%
- summarise(mean_rt = mean(rt), .groups = "drop") %>%
- mutate(mean_log_rt = log10(mean_rt))
- # Merge with sensorimotor norms (for Part 1)
- df_sm_acc <- df_rawc_with_norms %>%
- select(word, same, version, sensorimotor_distance,
- action_distance, perceptual_distance) %>%
- inner_join(df_tb2023_agg_acc, by = c("word", "same", "version"))
- df_sm_rt <- df_rawc_with_norms %>%
- select(word, same, version, sensorimotor_distance,
- action_distance, perceptual_distance) %>%
- inner_join(df_tb2023_agg_rt, by = c("word", "same", "version"))
- # Merge with full distributional sweep (for Part 2)
- df_full_acc <- df_merged %>%
- inner_join(df_tb2023_agg_acc, by = c("word", "same", "version"))
- df_full_rt <- df_merged %>%
- inner_join(df_tb2023_agg_rt, by = c("word", "same", "version"))
- # Sanity check
- nrow(df_sm_acc)
- nrow(df_sm_rt)
- nrow(df_full_acc)
- nrow(df_full_rt)
- ```
- ## Part 1: SM on its own
- ```{r}
- mod_sm_acc = lmer(data = df_sm_acc,
- accuracy ~ sensorimotor_distance +
- (1 | word))
- summary(mod_sm_acc)
- mod_sm_rt = lmer(data = df_sm_rt,
- mean_log_rt ~ sensorimotor_distance +
- (1 | word))
- summary(mod_sm_rt)
- ```
- ## Part 2: Comopare to distributional distance
- Here, we compare sensorimotor distance to the best-performing distributional distance measure from the RAW-C analysis.
- ```{r sm_dd_tb2023}
- # --- Compute AIC for each model/layer: SM, Distributional, Hybrid ---
- tb_acc_results <- df_full_acc %>%
- group_by(Model, Layer, n_params, Multilingual) %>%
- summarize(
- r2_sm = summary(lm(accuracy ~ sensorimotor_distance))$r.squared,
- aic_sm = AIC(lm(accuracy ~ sensorimotor_distance)),
- r2_bert = summary(lm(accuracy ~ Distance))$r.squared,
- aic_bert = AIC(lm(accuracy ~ Distance)),
- r2_combined = summary(lm(accuracy ~ Distance + sensorimotor_distance))$r.squared,
- aic_combined = AIC(lm(accuracy ~ Distance + sensorimotor_distance)),
- .groups = "drop"
- ) %>%
- mutate(
- aic_bert_vs_sm = aic_bert - aic_sm,
- aic_combined_vs_sm = aic_sm - aic_combined,
- aic_combined_vs_dist = aic_bert - aic_combined
- )
- tb_rt_results <- df_full_rt %>%
- group_by(Model, Layer, n_params, Multilingual) %>%
- summarize(
- r2_sm = summary(lm(mean_log_rt ~ sensorimotor_distance))$r.squared,
- aic_sm = AIC(lm(mean_log_rt ~ sensorimotor_distance)),
- r2_bert = summary(lm(mean_log_rt ~ Distance))$r.squared,
- aic_bert = AIC(lm(mean_log_rt ~ Distance)),
- r2_combined = summary(lm(mean_log_rt ~ Distance + sensorimotor_distance))$r.squared,
- aic_combined = AIC(lm(mean_log_rt ~ Distance + sensorimotor_distance)),
- .groups = "drop"
- ) %>%
- mutate(
- aic_bert_vs_sm = aic_bert - aic_sm,
- aic_combined_vs_sm = aic_sm - aic_combined,
- aic_combined_vs_dist = aic_bert - aic_combined
- )
- # --- Best layer per model (selected by best distributional R²) ---
- best_layers_acc <- tb_acc_results %>%
- group_by(Model, n_params, Multilingual) %>%
- slice_max(r2_bert, n = 1) %>%
- ungroup()
- best_layers_rt <- tb_rt_results %>%
- group_by(Model, n_params, Multilingual) %>%
- slice_max(r2_bert, n = 1) %>%
- ungroup()
- print(best_layers_acc)
- print(best_layers_rt)
- # --- Summary counts ---
- best_layers_acc %>%
- mutate(dist_better_than_sm = aic_bert_vs_sm < -4,
- hybrid_better_than_sm = aic_combined_vs_sm > 4,
- hybrid_better_than_dist = aic_combined_vs_dist > 4) %>%
- summarise(across(c(dist_better_than_sm,
- hybrid_better_than_sm,
- hybrid_better_than_dist), mean))
- best_layers_rt %>%
- mutate(dist_better_than_sm = aic_bert_vs_sm < -4,
- hybrid_better_than_sm = aic_combined_vs_sm > 4,
- hybrid_better_than_dist = aic_combined_vs_dist > 4) %>%
- summarise(across(c(dist_better_than_sm,
- hybrid_better_than_sm,
- hybrid_better_than_dist), mean))
- # --- LaTeX tables ---
- best_layers_acc %>%
- select(Model, aic_bert, aic_combined) %>%
- mutate(across(starts_with("aic"), round)) %>%
- kbl(col.names = c("Model", "Distributional", "Hybrid"),
- format = "latex", booktabs = TRUE,
- caption = sprintf("AIC values for predicting accuracy (Trott \\& Bergen, 2023). AIC for Sensorimotor Distance alone was %.2f.",
- unique(best_layers_acc$aic_sm)),
- label = "aic_tb_acc") %>%
- kable_styling(latex_options = "hold_position")
- best_layers_rt %>%
- select(Model, aic_bert, aic_combined) %>%
- mutate(across(starts_with("aic"), round)) %>%
- kbl(col.names = c("Model", "Distributional", "Hybrid"),
- format = "latex", booktabs = TRUE,
- caption = sprintf("AIC values for predicting log RT (Trott \\& Bergen, 2023). AIC for Sensorimotor Distance alone was %.2f.",
- unique(best_layers_rt$aic_sm)),
- label = "aic_tb_rt") %>%
- kable_styling(latex_options = "hold_position")
- # ============================================================
- # Figures: AIC dot plots (matching your existing style)
- # ============================================================
- # --- Accuracy ---
- aic_long_acc <- best_layers_acc %>%
- select(Model, Layer, aic_sm, aic_bert, aic_combined) %>%
- pivot_longer(cols = c(aic_sm, aic_bert, aic_combined),
- names_to = "type", values_to = "aic") %>%
- mutate(type = case_when(
- type == "aic_sm" ~ "Sensorimotor",
- type == "aic_bert" ~ "Distributional",
- type == "aic_combined" ~ "Hybrid"
- ),
- type = factor(type, levels = c("Sensorimotor", "Distributional", "Hybrid")))
- baseline_acc <- aic_long_acc %>% filter(type == "Sensorimotor") %>% pull(aic) %>% unique()
- aic_long_acc_no_sm <- aic_long_acc %>% filter(type != "Sensorimotor")
- aic_se_acc <- aic_long_acc_no_sm %>%
- group_by(type) %>%
- summarize(mean_aic = mean(aic),
- se_aic = sd(aic) / sqrt(n()),
- .groups = "drop")
- # Dynamic offset for the label (5% of y-range above baseline)
- offset_acc <- diff(range(aic_long_acc_no_sm$aic)) * 0.05
- ggplot() +
- geom_hline(yintercept = baseline_acc,
- linetype = "dashed", color = "coral", linewidth = 1) +
- geom_jitter(data = aic_long_acc_no_sm,
- aes(x = type, y = aic, color = type),
- alpha = 0.3, width = 0.1, size = 2) +
- geom_point(data = aic_se_acc,
- aes(x = type, y = mean_aic, color = type),
- size = 5) +
- geom_errorbar(data = aic_se_acc,
- aes(x = type, ymin = mean_aic - se_aic,
- ymax = mean_aic + se_aic, color = type),
- width = 0.15, linewidth = 1.2) +
- annotate("text", x = 1.5, y = baseline_acc + offset_acc,
- label = "Sensorimotor", color = "coral", size = 5) +
- labs(x = NULL, y = "AIC (lower is better)",
- title = "Predicting Accuracy") +
- scale_color_manual(values = c("Distributional" = "gray60",
- "Hybrid" = "steelblue")) +
- theme_minimal() +
- theme(text = element_text(size = 16), legend.position = "none")
- # --- RT ---
- aic_long_rt <- best_layers_rt %>%
- select(Model, Layer, aic_sm, aic_bert, aic_combined) %>%
- pivot_longer(cols = c(aic_sm, aic_bert, aic_combined),
- names_to = "type", values_to = "aic") %>%
- mutate(type = case_when(
- type == "aic_sm" ~ "Sensorimotor",
- type == "aic_bert" ~ "Distributional",
- type == "aic_combined" ~ "Hybrid"
- ),
- type = factor(type, levels = c("Sensorimotor", "Distributional", "Hybrid")))
- baseline_rt <- aic_long_rt %>% filter(type == "Sensorimotor") %>% pull(aic) %>% unique()
- aic_long_rt_no_sm <- aic_long_rt %>% filter(type != "Sensorimotor")
- aic_se_rt <- aic_long_rt_no_sm %>%
- group_by(type) %>%
- summarize(mean_aic = mean(aic),
- se_aic = sd(aic) / sqrt(n()),
- .groups = "drop")
- offset_rt <- diff(range(aic_long_rt_no_sm$aic)) * 0.05
- ggplot() +
- geom_hline(yintercept = baseline_rt,
- linetype = "dashed", color = "coral", linewidth = 1) +
- geom_jitter(data = aic_long_rt_no_sm,
- aes(x = type, y = aic, color = type),
- alpha = 0.3, width = 0.1, size = 2) +
- geom_point(data = aic_se_rt,
- aes(x = type, y = mean_aic, color = type),
- size = 5) +
- geom_errorbar(data = aic_se_rt,
- aes(x = type, ymin = mean_aic - se_aic,
- ymax = mean_aic + se_aic, color = type),
- width = 0.15, linewidth = 1.2) +
- annotate("text", x = 1.5, y = baseline_rt + offset_rt,
- label = "Sensorimotor", color = "coral", size = 5) +
- labs(x = NULL, y = "AIC (lower is better)",
- title = "Predicting RT") +
- scale_color_manual(values = c("Distributional" = "gray60",
- "Hybrid" = "steelblue")) +
- theme_minimal() +
- theme(text = element_text(size = 16), legend.position = "none")
- ```
- ## Part 3: Does SM account for sense boundaries?
- ```{r}
- # --- Compute AIC for each model/layer: same alone vs. same + continuous ---
- tb_acc_q3 <- df_full_acc %>%
- group_by(Model, Layer, n_params, Multilingual) %>%
- summarize(
- aic_same = AIC(lm(accuracy ~ same)),
- aic_same_sm = AIC(lm(accuracy ~ same + sensorimotor_distance)),
- aic_same_dist = AIC(lm(accuracy ~ same + Distance)),
- aic_same_both = AIC(lm(accuracy ~ same + Distance + sensorimotor_distance)),
- r2_same_dist = summary(lm(accuracy ~ same + Distance))$r.squared,
- .groups = "drop"
- ) %>%
- mutate(
- delta_aic_sm = aic_same - aic_same_sm, # positive = SM adds over same
- delta_aic_dist = aic_same - aic_same_dist, # positive = Distributional adds over same
- delta_aic_both = aic_same - aic_same_both # positive = both add over same
- )
- tb_rt_q3 <- df_full_rt %>%
- group_by(Model, Layer, n_params, Multilingual) %>%
- summarize(
- aic_same = AIC(lm(mean_log_rt ~ same)),
- aic_same_sm = AIC(lm(mean_log_rt ~ same + sensorimotor_distance)),
- aic_same_dist = AIC(lm(mean_log_rt ~ same + Distance)),
- aic_same_both = AIC(lm(mean_log_rt ~ same + Distance + sensorimotor_distance)),
- r2_same_dist = summary(lm(mean_log_rt ~ same + Distance))$r.squared,
- .groups = "drop"
- ) %>%
- mutate(
- delta_aic_sm = aic_same - aic_same_sm,
- delta_aic_dist = aic_same - aic_same_dist,
- delta_aic_both = aic_same - aic_same_both
- )
- # --- Best layer per model (selected by best same + Distributional R²) ---
- best_layers_acc_q3 <- tb_acc_q3 %>%
- group_by(Model, n_params, Multilingual) %>%
- slice_max(r2_same_dist, n = 1) %>%
- ungroup()
- best_layers_rt_q3 <- tb_rt_q3 %>%
- group_by(Model, n_params, Multilingual) %>%
- slice_max(r2_same_dist, n = 1) %>%
- ungroup()
- print(best_layers_acc_q3)
- print(best_layers_rt_q3)
- # --- Summary counts: how often does each augmentation add over same alone ---
- best_layers_acc_q3 %>%
- mutate(sm_adds = delta_aic_sm > 4,
- dist_adds = delta_aic_dist > 4,
- both_add = delta_aic_both > 4) %>%
- summarise(across(c(sm_adds, dist_adds, both_add), mean))
- best_layers_rt_q3 %>%
- mutate(sm_adds = delta_aic_sm > 4,
- dist_adds = delta_aic_dist > 4,
- both_add = delta_aic_both > 4) %>%
- summarise(across(c(sm_adds, dist_adds, both_add), mean))
- make_aic_plot_q3 <- function(best_layers_df, title) {
- aic_long <- best_layers_df %>%
- select(Model, Layer, aic_same, aic_same_sm, aic_same_dist, aic_same_both) %>%
- pivot_longer(cols = c(aic_same, aic_same_sm, aic_same_dist, aic_same_both),
- names_to = "type", values_to = "aic") %>%
- mutate(type = case_when(
- type == "aic_same" ~ "Same only",
- type == "aic_same_sm" ~ "Same + SM",
- type == "aic_same_dist" ~ "Same + Distributional",
- type == "aic_same_both" ~ "Same + Both"
- ),
- type = factor(type, levels = c("Same only", "Same + SM",
- "Same + Distributional", "Same + Both")))
- baseline <- aic_long %>% filter(type == "Same only") %>% pull(aic) %>% unique()
- aic_long_no_base <- aic_long %>% filter(type != "Same only")
- aic_se <- aic_long_no_base %>%
- group_by(type) %>%
- summarize(mean_aic = mean(aic),
- se_aic = sd(aic) / sqrt(n()),
- .groups = "drop")
- offset <- diff(range(aic_long_no_base$aic)) * 0.05
- ggplot() +
- geom_hline(yintercept = baseline,
- linetype = "dashed", color = "coral", linewidth = 1) +
- geom_jitter(data = aic_long_no_base,
- aes(x = type, y = aic, color = type),
- alpha = 0.3, width = 0.1, size = 2) +
- geom_point(data = aic_se,
- aes(x = type, y = mean_aic, color = type),
- size = 5) +
- geom_errorbar(data = aic_se,
- aes(x = type, ymin = mean_aic - se_aic,
- ymax = mean_aic + se_aic, color = type),
- width = 0.15, linewidth = 1.2) +
- annotate("text", x = 2, y = baseline + offset,
- label = "Same only", color = "coral", size = 5) +
- labs(x = NULL, y = "AIC (lower is better)", title = title) +
- scale_color_manual(values = c("Same + SM" = "coral",
- "Same + Distributional" = "gray60",
- "Same + Both" = "steelblue")) +
- theme_minimal() +
- theme(text = element_text(size = 16), legend.position = "none",
- axis.text.x = element_text(angle = 20, hjust = 1))
- }
- make_aic_plot_q3(best_layers_acc_q3, "Predicting Accuracy")
- make_aic_plot_q3(best_layers_rt_q3, "Predicting RT")
- ```
- # Comparison to Trott (2024) LLM Norms
- Here, we conduct a more in-depth comparison to the GPT-4-generated norms from Trott (2024).
- ## Load data
- ```{r}
- df_gpt_perception = read_csv("../../data/trott2024/cs_norms_perception_gpt-4.csv")
- df_gpt_action = read_csv("../../data/trott2024/cs_norms_action_gpt-4.csv")
- df_both_gpt = df_gpt_perception %>%
- inner_join(df_gpt_action)
- df_gpt_with_human = df_both_gpt %>%
- inner_join(df_contextualized_meanings)
- ### RAW-C With GPT and human SM distance
- df_rawc_with_sm = read_csv("../../data/trott2024/rawc_with_sm.csv")
- ```
- ## Dimension-level correlation
- ```{r gpt_corrs}
- df_summ = df_gpt_with_human %>%
- summarise(Vision = cor(Vision.M, Vision, method = "spearman"),
- Hearing = cor(Hearing.M, Hearing, method = "spearman"),
- Touch = cor(Touch.M, Touch, method = "spearman"),
- Olfaction = cor(Olfaction.M, Olfaction, method = "spearman"),
- Taste = cor(Taste.M, Taste, method = "spearman"),
- Interoception = cor(Interoception.M, Interoception, method = "spearman"),
- Hand_arm = cor(Hand_arm.M, Hand_arm, method = "spearman"),
- Foot_leg = cor(Foot_leg.M, Foot_leg, method = "spearman"),
- Head = cor(Head.M, Head, method = "spearman"),
- Torso = cor(Torso.M, Torso, method = "spearman"),
- Mouth_throat = cor(Mouth_throat.M, Mouth_throat, method = "spearman"))
- df_summ
- df_long = df_summ %>%
- pivot_longer(everything(), names_to = "Factor", values_to = "Correlation")
- df_long %>%
- ggplot(aes(x = reorder(Factor, Correlation), y = Correlation)) +
- geom_bar(stat = "identity") +
- labs(x = "Factor", y = "Correlation") +
- theme_minimal()
- ```
- ## Compare correlation matrices
- ```{r gpt_corrs_comparison}
- library(tidyverse)
- library(ggcorrplot)
- library(patchwork)
- dim_order <- c("Vision", "Hearing", "Olfaction", "Taste", "Interoception",
- "Touch", "Mouth_throat", "Head", "Torso", "Hand_arm", "Foot_leg")
- # Human matrix: .M columns
- human_cols <- df_gpt_with_human %>%
- select(all_of(paste0(dim_order, ".M"))) %>%
- rename_with(~ str_remove(., "\\.M"))
- cors_human <- cor(human_cols, use = "pairwise.complete.obs")
- # GPT matrix: bare-named columns
- gpt_cols <- df_gpt_with_human %>%
- select(all_of(dim_order))
- cors_gpt <- cor(gpt_cols, use = "pairwise.complete.obs")
- # Difference (GPT - Human)
- cors_diff <- cors_gpt - cors_human
- ggcorrplot(cors_human, hc.order = FALSE, type = "upper",
- lab = TRUE, lab_size = 2.5) +
- ggtitle("Human") +
- theme(axis.text.x = element_text(size = 10, angle = 45, hjust = 1))
- ggcorrplot(cors_gpt, hc.order = FALSE, type = "upper",
- lab = TRUE, lab_size = 2.5) +
- ggtitle("GPT-4") +
- theme(axis.text.x = element_text(size = 10, angle = 45, hjust = 1))
- ggcorrplot(cors_diff, hc.order = FALSE, type = "upper",
- lab = TRUE, lab_size = 2.5,
- colors = c("#2166AC", "white", "#B2182B")) +
- ggtitle("GPT − Human") +
- theme(axis.text.x = element_text(size = 10, angle = 45, hjust = 1))
- ### Statistical comparison
- library(vegan)
- # Mantel test
- mantel_result <- mantel(cors_human, cors_gpt,
- method = "pearson", # or "spearman"
- permutations = 9999)
- mantel_result
- mantel_result <- mantel(cors_human, cors_gpt,
- method = "spearman", # or "spearman"
- permutations = 9999)
- mantel_result
- ## Or just straight-up correlation
- upper_human <- cors_human[upper.tri(cors_human)]
- upper_gpt <- cors_gpt[upper.tri(cors_gpt)]
- # Correlation between the two
- cor(upper_human, upper_gpt, method = "pearson")
- cor(upper_human, upper_gpt, method = "spearman")
- ```
- ## Substitution analyses
- ### Homonymy vs. Polysemy
- Here, we ask whether homonyms show bigger sensorimotor distance for GPT-generated norms as well.
- ```{r}
- model_full = lmer(data = df_rawc_with_sm,
- sm_gpt ~ same +
- (1 + same | word),
- control=lmerControl(optimizer="bobyqa"),
- REML = FALSE)
- model_reduced = lmer(data = df_rawc_with_sm,
- sm_gpt ~ # same +
- (1 + same | word),
- control=lmerControl(optimizer="bobyqa"),
- REML = FALSE)
- summary(model_full)
- anova(model_full, model_reduced)
- model_full_at = lmer(data = df_rawc_with_sm,
- sm_gpt ~ same * ambiguity_type +
- (1 + same | word),
- control=lmerControl(optimizer="bobyqa"),
- REML = FALSE)
- model_reduced_at = lmer(data = df_rawc_with_sm,
- sm_gpt ~ same + ambiguity_type +
- (1 + same | word),
- control=lmerControl(optimizer="bobyqa"),
- REML = FALSE)
- summary(model_full_at)
- anova(model_full_at, model_reduced_at)
- df_rawc_with_sm %>%
- group_by(same, ambiguity_type) %>%
- summarise(mean_gpt = mean(sm_gpt),
- sd_gpt = sd(sm_gpt),
- mean_human = mean(sm_human),
- sd_human = sd(sm_human))
- df_rawc_with_sm = df_rawc_with_sm %>%
- mutate(Same = case_when(
- same == TRUE ~ "Same Sense",
- same == FALSE ~ "Different Sense"
- ))
- df_rawc_with_sm %>%
- ggplot(aes(x = sm_gpt,
- y = ambiguity_type,
- fill = Same)) +
- geom_density_ridges2(aes(height = ..density..),
- color = NA,
- alpha = 0.5,
- scale=0.85,
- stat="density") +
- labs(x = "Sensorimotor Distance (GPT-4)",
- y = NULL,
- fill = NULL) +
- scale_fill_manual(
- values = c("Same Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[1],
- "Different Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[2])
- ) +
- theme_minimal() +
- theme(text = element_text(size = 20))
- ```
- ### Dominance
- We also replicate the dominance analysis.
- ```{r gpt_sm_strength}
- df_dom_gpt = df_gpt_with_human %>%
- inner_join(df_dominance_individual)
- df_dom_gpt <- df_dom_gpt %>%
- inner_join(df_lancaster)
- nrow(df_dom_gpt)
- df_dom_gpt = df_dom_gpt %>%
- rowwise() %>%
- mutate(max_strength_gpt = max(
- c(
- ## Modalities
- Vision,
- Hearing,
- Olfaction,
- Touch,
- Taste,
- Interoception,
- ## Effectors
- Head,
- Mouth_throat,
- Torso,
- Hand_arm,
- Foot_leg
- )
- ),
- minkowski3_strength_gpt = sum(c(Vision,
- Hearing,
- Olfaction,
- Touch,
- Taste,
- Interoception,
- ## Effectors
- Head,
- Mouth_throat,
- Torso,
- Hand_arm,
- Foot_leg)^3)^(1/3),
- max_strength = max(
- c(
- ## Modalities
- Vision.M,
- Hearing.M,
- Olfaction.M,
- Touch.M,
- Taste.M,
- Interoception.M,
- ## Effectors
- Head.M,
- Mouth_throat.M,
- Torso.M,
- Hand_arm.M,
- Foot_leg.M
- )
- ),
- minkowski3_strength = sum(c(Vision.M, Hearing.M, Olfaction.M, Touch.M, Taste.M,
- Interoception.M, Head.M, Mouth_throat.M, Torso.M,
- Hand_arm.M, Foot_leg.M)^3)^(1/3)
- ) %>%
- ungroup()
- cor.test(df_dom_gpt$max_strength_gpt, df_dom_gpt$max_strength)
- cor.test(df_dom_gpt$minkowski3_strength_gpt, df_dom_gpt$minkowski3_strength)
- df_dom_gpt %>%
- ggplot(aes(x = minkowski3_strength_gpt,
- y = minkowski3_strength)) +
- geom_point(alpha = .5) +
- geom_smooth(method = "lm") +
- labs(x = "Minkowski-3 Strength (GPT-4)",
- y = "Minkowski-3 Strength (Human)") +
- theme_minimal() +
- theme(text = element_text(size = 16), legend.position = "none")
- ```
- ### Does GPT sensorimotor strength predict dominance?
- ```{r gpt_dominance}
- df_dom_gpt %>%
- ggplot(aes(x = minkowski3_strength,
- y = dominance)) +
- geom_point(alpha = .4) +
- geom_smooth(method = "lm") +
- labs(x = "Minkowski-3 Sensorimotor Strength (GPT-4)",
- y = "Dominance") +
- theme_minimal() +
- theme(text = element_text(size=16))
- model_full = lmer(data = df_dom_gpt,
- dominance ~
- max_strength_gpt +
- Max_strength.sensorimotor + Minkowski3.sensorimotor +
- (1 | word),
- REML = FALSE)
- model_reduced = lmer(data = df_dom_gpt,
- dominance ~
- # max_strength_gpt +
- Max_strength.sensorimotor + Minkowski3.sensorimotor +
- (1 | word),
- REML = FALSE)
- summary(model_full)
- anova(model_full, model_reduced)
- model_full = lmer(data = df_dom_gpt,
- dominance ~
- minkowski3_strength_gpt +
- Max_strength.sensorimotor + Minkowski3.sensorimotor +
- (1 | word),
- REML = FALSE)
- model_reduced = lmer(data = df_dom_gpt,
- dominance ~
- # minkowski3_strength_gpt +
- Max_strength.sensorimotor + Minkowski3.sensorimotor +
- (1 | word),
- REML = FALSE)
- summary(model_full)
- anova(model_full, model_reduced)
- ```
contextualized_norms_analysis.Rmd at commit 7f87ad2, no license · at the source
Overview
Abstract
Embodied theories of language emphasize the role of sensorimotor experience in linguistic knowledge. Central to testing these theories is the creation of large datasets of linguistic norms, which contain judgments about a word’s sensorimotor associations and can be used to predict human behavioral or brain data – sometimes in contrast to competing variables, such as those derived from distributional language models. Yet many of these datasets contain judgments about words in isolation, despite the fact that most words are ambiguous, making it difficult to determine which meaning of a word is characterized by its rating (e.g., “wooden table” vs. “data table”). In the current work, we introduce a new lexical resource (directly inspired by the Lancaster sensorimotor for 112 English words, each rated in four different contexts (448 sentences total). We demonstrate: first, that these ratings encode overlapping but distinct information from the Lancaster sensorimotor norms; second, that decontextualized ratings likely reflect the more dominant meaning of ambiguous words; third, that homonyms have more distinct sensorimotor profiles than polysemes; fourth, that the contextualized sensorimotor distance between two uses of an ambiguous word predicts human judgments about semantic relatedness; and fifth, that ratings derived from GPT-4 align reasonably well with human judgments. We conclude by suggesting that contextualized ratings like these can be used both to inform competing theories of semantic representations and also to evaluate or “probe” the ability of LMs to recover sensorimotor information.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 6 matches between paragraphs and lines of code.
seantrott/cs_norms
7f87ad29ff1280087f995f7e691a4bb0ab12ffcb, 1 June 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
4 files
- src/
analysis/ , R, 2,607 lines, 6 matchescontextualized_norms_ana lysis.Rmd - src/
data/ , Python, 76 linesbow.py - src/
data/ , Python, 45 linesdistance_to_ls.py - README.md, Text, 43 lines
Code Availability
Analysis code can be found on GitHub: https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 3 scripts, each with its path and the digest of its content;
- 6 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- huggingface.co/
datasets/ , at Hugging Face; found in “Data Availability”seantrott/ cs_norms
Data Availability Statement
The contextualized sensorimotor norms are publicly available on GitHub (https://
Analysis code can be found on GitHub: https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Open Practices Statement
All code and data can be found in a publicly available repository (see above).
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 3, 28 September 2026
- Publisher: n/a → Springer Science+Business Media
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 2 authors, 7 MeSH terms, 59 references.
Cite
This paper
Trott, S., & Bergen, B. (2026). Contextualized sensorimotor norms: Multi-dimensional measures of sensorimotor strength for ambiguous English words, in context. Behavior research methods, 58(9), 264. https://
BibTeX
@article{trott2026contex
author = {Trott, Sean and Bergen, Benjamin},
title = {{Contextualized sensorimotor norms: Multi-dimensional measures of sensorimotor strength for ambiguous English words, in context}},
journal = {Behavior research methods},
year = {2026},
month = aug,
volume = {58},
number = {9},
pages = {264},
publisher = {Springer Science+Business Media},
issn = {1554-351X},
doi = {10.3758/
url = {https://
pmid = {42587187},
pmcid = {PMC13469423}
}
RIS
TY - JOUR
AU - Trott, Sean
AU - Bergen, Benjamin
TI - Contextualized sensorimotor norms: Multi-dimensional measures of sensorimotor strength for ambiguous English words, in context
T2 - Behavior research methods
J2 - Behav Res Methods
PY - 2026
DA - 2026/
VL - 58
IS - 9
SP - 264
SN - 1554-351X
PB - Springer Science+Business Media
DO - 10.3758/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.3758/
"type": "article-journal",
"title": "Contextualized sensorimotor norms: Multi-dimensional measures of sensorimotor strength for ambiguous English words, in context",
"container-title": "Behavior research methods",
"author": [
{
"family": "Trott",
"given": "Sean"
},
{
"family": "Bergen",
"given": "Benjamin"
}
],
"container-title-short":
"volume": "58",
"issue": "9",
"page": "264",
"DOI": "10.3758/
"PMID": "42587187",
"PMCID": "PMC13469423",
"ISSN": "1554-351X",
"publisher": "Springer Science+Business Media",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
12
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1073/pnas.2603114123 [code]
- The human hippocampus can pattern separate memories by meaning.Journal: Proceedings of the National Academy of Sciences of the United States of AmericaIn common: broom, lmerTest, lme4, 5 other tools
- [2] doi:10.1016/j.dcn.2026.101765 [code]
- Fusiform face area development correlates with development in higher-order social brain regions.Journal: Developmental cognitive neuroscienceIn common: broom, lmerTest, lme4, 5 other tools
- [3] doi:10.1162/imag.a.1235 [code]
- Intracranial volume: To adjust or not to adjust? It is not a matter of if, but how.Journal: Imaging neuroscience (Cambridge, Mass.)In common: broom, lmerTest, lme4, 5 other tools
- [4] doi:10.64898/2026.03.09.710596 [code]
- Infant gut microbiomes contribute to metabolic states that impact brain functionJournal: bioRxiv (preprint)In common: broom, lmerTest, lme4, 4 other tools, 1 reference
- [5] doi:10.1162/imag.a.1321 [code]
- Phase similarity between similar objects indicates representational merging across retrieval training but not sleep.Journal: Imaging neuroscience (Cambridge, Mass.)In common: broom, lmerTest, lme4, 4 other tools, 1 reference
- [6] doi:10.1093/cercor/bhag132 [code]
- Spatiotemporal white-matter development across early childhood.Journal: Cerebral cortex (New York, N.Y. : 1991)In common: broom, lmerTest, lme4, 4 other tools
- [7] doi:10.1038/s41467-026-71264-8 [code]
- Habitual coffee intake shapes the gut microbiome and modifies host physiology and cognition.Journal: Nature communicationsIn common: broom, lmerTest, lme4, 4 other tools
- [8] doi:10.1016/j.celrep.2026.117505 [code]
- Impaired spatial coding and neuronal hyperactivity in the medial entorhinal cortex of aged APP knock-in mice.Journal: Cell reportsIn common: broom, lmerTest, lme4, 4 other tools
- [9] doi:10.3390/ijms27135713 [code]
- Chronic Administration of Marinobufagenin in Mice Causes Hyperlocomotion and Decrease in Anxiety by Altering Monoamine Turnover Unaccompanied by Motor Deficits or Oxidative Stress.Journal: International journal of molecular sciencesIn common: broom, lmerTest, lme4, 4 other tools
- [10] doi:10.1093/ageing/afag263 [code]
- Cardiometabolic medication exposures and cognitive outcomes in Alzheimer's disease.Journal: Age and ageingIn common: broom, lmerTest, lme4, 4 other tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 3 scripts, and 6 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:67e36fac0c26a73e…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
