OSCR

Contextualized sensorimotor norms: Multi-dimensional measures of sensorimotor strength for ambiguous English words, in context.

Code ↔ Paper

6 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 6 matches
  1. [1] § Case study: Can large language models help scale the CS norms? › Between-dimension comparison ↔ src/analysis/contextualized_norms_analysis.Rmd, lines 2313–2381 · score 0.67 · Mantel, permutations, pairwise, Pearson, matrices, Spearman
  2. [2] § Contextual variance in sensorimotor associations › Sense dominance and deviation from the Lancaster norms ↔ src/analysis/contextualized_norms_analysis.Rmd, lines 1170–1200 · score 0.66 · decontextualized LS norm, cosine distance, sensorimotor properties, lme4, dominant sense, closer
  3. [3] § Comparing distributional and sensorimotor similarity › Methods › Calculating distributional distance ↔ src/analysis/contextualized_norms_analysis.Rmd, lines 1254–1326 · score 0.65 · DistilBERT, RoBERTa, ALBERT, RAW, distance, models
  4. [4] § Contextual variance in sensorimotor associations › Deviations from the Lancaster sensorimotor norms ↔ src/analysis/contextualized_norms_analysis.Rmd, lines 766–891 · score 0.64 · standardized absolute deviation, maximally, floor, Olfaction, SD, Taste
  5. [5] § Comparing distributional and sensorimotor similarity › Sensorimotor distance and ambiguity processing › Explanation of experimental materials ↔ src/analysis/contextualized_norms_analysis.Rmd, lines 1870–1918 · score 0.62 · primed sensibility judgment, Log RT, aggregated, accuracy, asked, word
  6. [6] § Comparing distributional and sensorimotor similarity › Sensorimotor distance and ambiguity processing › Evaluating sensorimotor vs. distributional distance ↔ src/analysis/contextualized_norms_analysis.Rmd, lines 1870–1918 · score 0.56 · primed sensibility judgment, Log RT, accuracy, sensorimotor distance

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R Markdown · 2,607 lines · 81 KB · no license · 6 matches

  1. ---
  2. title: "Analysis of contextualized sensorimotor norms"
  3. author: "Sean Trott"
  4. date: "January 19, 2026"
  5. output:
  6. html_document:
  7. keep_md: yes
  8. toc: yes
  9. toc_float: yes
  10. ---
  11. ```{r setup, include=FALSE}
  12. knitr::opts_chunk$set(dpi = 300, fig.format = "png")
  13. ```
  14. ```{r include=FALSE}
  15. library(tidyverse)
  16. library(lme4)
  17. library(ggridges)
  18. library(broom.mixed)
  19. library(lmerTest)
  20. library(ggcorrplot)
  21. library(kableExtra)
  22. ```
  23. # Introduction
  24. In this document, we analyze the **contextualized sensorimotor norms**: judgments about the strength of different sensorimotor dimensions of ambiguous words, in context.
  25. We use these norms in several analyses:
  26. 1. First, we compare them to the corresponding dimensions for the "decontextualized" Lancaster sensorimotor norms.
  27. 2. Second, we ask whether the **dominance** of a word sense is correlated with its sensorimotor strength, i.e., whether more concrete meanings tend to be rated as more dominant.
  28. 3. Third, we ask whether the **sensorimotor distance** between two contexts of use predicts judgments of how **related** those meanings are, above and beyond their distributional similarity and whether or not they belong to the same sense.
  29. 4. Fourth, we use **sensorimotor distance** to predict behavior on a primed sensibility judgment task.
  30. # Characterizing dimensions
  31. First, load the data.
  32. ```{r}
  33. # setwd("/Users/seantrott/Dropbox/UCSD/Research/Ambiguity/SSD/cs_norms/src/analysis")
  34. df_contextualized_meanings = read_csv("../../data/processed/contextualized_sensorimotor_norms.csv")
  35. nrow(df_contextualized_meanings)
  36. ```
  37. ## Visualizing distributions
  38. ```{r different_dimensions}
  39. df_contextualized_meanings_long = df_contextualized_meanings %>%
  40. pivot_longer(cols = c(Vision.M, Hearing.M, Olfaction.M,
  41. Taste.M, Interoception.M, Touch.M,
  42. Mouth_throat.M, Head.M, Torso.M,
  43. Hand_arm.M, Foot_leg.M),
  44. names_to = "Dimension",
  45. values_to = "Strength") %>%
  46. mutate(Dimension = gsub('.M', '', Dimension))
  47. df_contextualized_meanings_long %>%
  48. ggplot(aes(x = reorder(Dimension, Strength),
  49. y = Strength)) +
  50. geom_violin() +
  51. geom_jitter(alpha = .1,
  52. width = .1) +
  53. coord_flip() +
  54. labs(y = "Sensorimotor strength",
  55. x = "Dimension") +
  56. theme_bw() +
  57. theme(text = element_text(size=20))
  58. ```
  59. ## Correlations across dimensions
  60. ```{r corr_matrix}
  61. columns = df_contextualized_meanings %>%
  62. mutate(Vision = Vision.M,
  63. Hearing = Hearing.M,
  64. Olfaction = Olfaction.M,
  65. Taste = Taste.M,
  66. Interoception = Interoception.M,
  67. Touch = Touch.M,
  68. Mouth_throat = Mouth_throat.M,
  69. Head = Head.M,
  70. Torso = Torso.M,
  71. Hand_arm = Hand_arm.M,
  72. Foot_leg = Foot_leg.M) %>%
  73. select(Vision, Hearing, Olfaction,
  74. Taste, Interoception, Touch,
  75. Mouth_throat, Head, Torso,
  76. Hand_arm, Foot_leg)
  77. cors = cor(columns)
  78. # cors[lower.tri(cors, diag=TRUE)] <- 0
  79. # Plot the correlation matrix
  80. ggcorrplot(cors,
  81. hc.order = FALSE,
  82. # method = "square",
  83. type = "upper") +
  84. theme(
  85. axis.text.x = element_text(size = 10, angle = 45, hjust = 1),
  86. axis.text.y = element_text(size = 10)
  87. )
  88. ggcorrplot(cors, hc.order = FALSE, type = "upper",
  89. lab = TRUE, lab_size = 2.5) +
  90. theme(axis.text.x = element_text(size = 10, angle = 45, hjust = 1))
  91. ```
  92. # Predicting dominance
  93. ## Load and merge data
  94. Load the item-level means for the sensorimotor norms.
  95. ```{r}
  96. df_contextualized_meanings = read_csv("../../data/processed/contextualized_sensorimotor_norms_with_ls.csv")
  97. nrow(df_contextualized_meanings)
  98. ```
  99. Load the dominance norms.
  100. ```{r}
  101. df_dominance = read_csv("../../data/processed/dominance_norms_with_order.csv")
  102. ## Determine the specific sense/meaning of the righthand context
  103. df_dominance = df_dominance %>%
  104. mutate(context = substr(version_with_order, 6, 9))
  105. ## Now group by that righthand context to get relative dominance of that meaning
  106. df_dominance_individual = df_dominance %>%
  107. group_by(word, context) %>%
  108. summarise(dominance = mean(dominance_right))
  109. nrow(df_dominance_individual)
  110. ```
  111. Merge the dominance and sensorimotor norms data.
  112. ```{r}
  113. df_dom_plus_sm = df_contextualized_meanings %>%
  114. inner_join(df_dominance_individual)
  115. nrow(df_dom_plus_sm)
  116. ```
  117. We also load and merge the Lancaster norms, as a control.
  118. ```{r}
  119. df_lancaster <- read_csv("../../data/lexical/lancaster_norms.csv") %>%
  120. mutate(word = tolower(Word)) %>%
  121. rename(
  122. Foot_leg.SD.LS = Foot_leg.SD,
  123. Hand_arm.SD.LS = Hand_arm.SD,
  124. Head.SD.LS = Head.SD,
  125. Torso.SD.LS = Torso.SD
  126. )
  127. df_dom_plus_sm <- df_dom_plus_sm %>%
  128. inner_join(df_lancaster)
  129. nrow(df_dom_plus_sm)
  130. ```
  131. ## Calculating contextualized sensorimotor strength
  132. Based on Lynott et al (2019), we **contextualized sensorimotor strength** as the *maximum* strength across all the dimensions.
  133. ```{r sm_strength_overall}
  134. df_dom_plus_sm = df_dom_plus_sm %>%
  135. rowwise() %>%
  136. mutate(max_strength = max(
  137. c(
  138. ## Modalities
  139. Vision.M,
  140. Hearing.M,
  141. Olfaction.M,
  142. Touch.M,
  143. Taste.M,
  144. Interoception.M,
  145. ## Effectors
  146. Head.M,
  147. Mouth_throat.M,
  148. Torso.M,
  149. Hand_arm.M,
  150. Foot_leg.M
  151. )
  152. ),
  153. minkowski3_strength = sum(c(Vision.M, Hearing.M, Olfaction.M, Touch.M, Taste.M,
  154. Interoception.M, Head.M, Mouth_throat.M, Torso.M,
  155. Hand_arm.M, Foot_leg.M)^3)^(1/3),
  156. max_perceptual_strength = max(
  157. c(
  158. ## Modalities
  159. Vision.M,
  160. Hearing.M,
  161. Olfaction.M,
  162. Touch.M,
  163. Taste.M,
  164. Interoception.M
  165. )
  166. ),
  167. max_action_strength = max(
  168. c(
  169. ## Effectors
  170. Head.M,
  171. Mouth_throat.M,
  172. Torso.M,
  173. Hand_arm.M,
  174. Foot_leg.M
  175. )
  176. )
  177. ) %>%
  178. ungroup()
  179. df_dom_plus_sm %>%
  180. ggplot(aes(x = Max_strength.sensorimotor,
  181. y = max_strength)) +
  182. geom_point(alpha = .5) +
  183. labs(y = "Maximum Contextualized Strength",
  184. x = "Maximum Strength (Lancaster)") +
  185. theme_bw()
  186. cor.test(df_dom_plus_sm$Max_strength.sensorimotor,
  187. df_dom_plus_sm$max_strength)
  188. cor.test(df_dom_plus_sm$max_strength,
  189. df_dom_plus_sm$minkowski3_strength)
  190. ```
  191. ## Does contextualized sensorimotor strength predict dominance?
  192. The answer is **yes**: contexts with a higher *maximum* sensorimotor strength also tend to be rated as more *dominant*.
  193. Notably, this is true above and beyond the *decontextualized* ratings of sensorimotor strength for a given word.
  194. ```{r strength_dominance}
  195. df_dom_plus_sm %>%
  196. ggplot(aes(x = max_strength,
  197. y = dominance)) +
  198. geom_point(alpha = .4) +
  199. geom_smooth(method = "lm") +
  200. labs(x = "Maximum Contextualized Strength",
  201. y = "Dominance") +
  202. theme_minimal() +
  203. theme(text = element_text(size=16))
  204. df_dom_plus_sm %>%
  205. ggplot(aes(x = minkowski3_strength,
  206. y = dominance)) +
  207. geom_point(alpha = .4) +
  208. geom_smooth(method = "lm") +
  209. labs(x = "Minkowski-3 Sensorimotor Strength",
  210. y = "Dominance") +
  211. theme_minimal() +
  212. theme(text = element_text(size=16))
  213. model_full = lmer(data = df_dom_plus_sm,
  214. dominance ~
  215. max_strength +
  216. Max_strength.sensorimotor + Minkowski3.sensorimotor +
  217. (1 | word),
  218. REML = FALSE)
  219. model_reduced = lmer(data = df_dom_plus_sm,
  220. dominance ~
  221. # max_strength +
  222. Max_strength.sensorimotor + Minkowski3.sensorimotor +
  223. (1 | word),
  224. REML = FALSE)
  225. summary(model_full)
  226. anova(model_full, model_reduced)
  227. model_full = lmer(data = df_dom_plus_sm,
  228. dominance ~
  229. minkowski3_strength +
  230. Max_strength.sensorimotor + Minkowski3.sensorimotor +
  231. (1 | word),
  232. REML = FALSE)
  233. model_reduced = lmer(data = df_dom_plus_sm,
  234. dominance ~
  235. # max_strength +
  236. Max_strength.sensorimotor + Minkowski3.sensorimotor +
  237. (1 | word),
  238. REML = FALSE)
  239. summary(model_full)
  240. anova(model_full, model_reduced)
  241. ```
  242. ### Examples
  243. ```{r}
  244. ### Top 10
  245. df_dom_plus_sm %>%
  246. arrange(desc(minkowski3_strength)) %>%
  247. select(word, sentence, minkowski3_strength, dominance) %>%
  248. head(10)
  249. ### Bottom 10
  250. df_dom_plus_sm %>%
  251. arrange(minkowski3_strength) %>%
  252. select(word, sentence, minkowski3_strength, dominance) %>%
  253. head(10)
  254. ### Top 10
  255. df_dom_plus_sm %>%
  256. arrange(desc(max_strength)) %>%
  257. select(word, sentence, max_strength, dominance) %>%
  258. head(10)
  259. ### Bottom 10
  260. df_dom_plus_sm %>%
  261. arrange(max_strength) %>%
  262. select(word, sentence, max_strength, dominance) %>%
  263. head(10)
  264. ```
  265. ### Which dimension best predicts dominance?
  266. Here, we ask whether specific dimensions are particularly correlated with sense dominance.
  267. ```{r specific_dimensions_dominance}
  268. features <- c("Vision.M", "Hearing.M", "Olfaction.M", "Touch.M", "Taste.M",
  269. "Interoception.M", "Head.M", "Mouth_throat.M", "Torso.M",
  270. "Hand_arm.M", "Foot_leg.M")
  271. # Models with each feature + baseline covariate
  272. r2_results <- map_dfr(features, function(feat) {
  273. formula <- as.formula(paste0("dominance ~ Max_strength.sensorimotor + ", feat))
  274. model <- lm(formula, data = df_dom_plus_sm)
  275. tibble(
  276. feature = feat,
  277. R2 = summary(model)$r.squared,
  278. beta = coef(model)[3], # coefficient for the feature (3rd term now)
  279. p_value = summary(model)$coefficients[3, 4]
  280. )
  281. })
  282. # Baseline model (just Max_strength.sensorimotor)
  283. baseline_model <- lm(dominance ~ Max_strength.sensorimotor, data = df_dom_plus_sm)
  284. baseline_row <- tibble(
  285. feature = "Baseline",
  286. R2 = summary(baseline_model)$r.squared,
  287. beta = NA,
  288. p_value = NA
  289. )
  290. # Combine
  291. r2_results_with_baseline <- bind_rows(baseline_row, r2_results)
  292. # Plot
  293. r2_results_with_baseline %>%
  294. mutate(feature = str_remove(feature, "\\.M$"),
  295. feature = str_replace(feature, "_", "/"),
  296. sig = ifelse(p_value < .05, "*", ""),
  297. is_baseline = feature == "Baseline",
  298. feature = fct_reorder(feature, R2)) %>%
  299. ggplot(aes(x = R2, y = feature, fill = is_baseline)) +
  300. geom_col() +
  301. geom_text(aes(label = sig), hjust = -0.5, size = 6, na.rm = TRUE) +
  302. scale_fill_manual(values = c("grey40", "steelblue"), guide = "none") +
  303. labs(x = "R²", y = NULL, title = "Predicting Sense Dominance") +
  304. theme_minimal() +
  305. theme(text = element_text(size = 16))
  306. cor.test(df_dom_plus_sm$Touch.M,
  307. df_dom_plus_sm$dominance)
  308. cor.test(df_dom_plus_sm$Vision.M,
  309. df_dom_plus_sm$dominance)
  310. ```
  311. It looks like vision and touch are especially strong predictors. Which items drive this, e.g., for touch?
  312. ```{r touch_dominance}
  313. by_sense <- df_dom_plus_sm %>%
  314. mutate(sense = str_extract(context, "M[12]")) %>%
  315. group_by(word, sense) %>%
  316. summarize(
  317. mean_touch = mean(Touch.M),
  318. mean_vision = mean(Vision.M),
  319. mean_dom = mean(dominance),
  320. .groups = "drop"
  321. )
  322. # Now 2 rows per word—compute within-word difference
  323. sense_diffs <- by_sense %>%
  324. pivot_wider(names_from = sense,
  325. values_from = c(mean_touch, mean_vision, mean_dom)) %>%
  326. mutate(
  327. touch_diff = mean_touch_M1 - mean_touch_M2,
  328. vision_diff = mean_vision_M1 - mean_vision_M2,
  329. dom_diff = mean_dom_M1 - mean_dom_M2
  330. )
  331. sense_diffs %>%
  332. mutate(abs_touch_diff = abs(touch_diff)) %>%
  333. arrange(desc(abs_touch_diff)) %>%
  334. select(word, abs_touch_diff, dom_diff) %>%
  335. head(5)
  336. sense_diffs %>%
  337. mutate(abs_vision_diff = abs(vision_diff)) %>%
  338. arrange(desc(abs_vision_diff)) %>%
  339. select(word, abs_vision_diff, dom_diff) %>%
  340. head(5)
  341. # Does the sense with higher Touch or Vision tend to be more dominant?
  342. cor.test(sense_diffs$touch_diff, sense_diffs$dom_diff)
  343. cor.test(sense_diffs$vision_diff, sense_diffs$dom_diff)
  344. ```
  345. # How much does context add?
  346. Here, we inspect how much information the CS norms provide, both relative to the LS norms and just overall (i.e., how much context pushes sensorimotor vectors around).
  347. ## Comparing dimensions to Lancaster
  348. Now, we compare each dimension to the LS Norms.
  349. ```{r ls_deviation_market}
  350. df_diffs_raw <- df_dom_plus_sm %>%
  351. mutate(vision_diff = (Vision.M - Visual.mean),
  352. auditory_diff = (Hearing.M - Auditory.mean),
  353. intero_diff = (Interoception.M - Interoceptive.mean),
  354. olfactory_diff = (Olfaction.M - Olfactory.mean),
  355. touch_diff = (Touch.M - Haptic.mean),
  356. taste_diff = (Taste.M - Gustatory.mean),
  357. torso_diff = (Torso.M - Torso.mean),
  358. hand_arm_diff = (Hand_arm.M - Hand_arm.mean),
  359. foot_leg_diff = (Foot_leg.M - Foot_leg.mean),
  360. head_diff = (Head.M - Head.mean),
  361. mouth_throat_diff = (Mouth_throat.M - Mouth.mean)) %>%
  362. pivot_longer(cols = ends_with("_diff"),
  363. names_to = "Dimension", values_to = "Diff") %>%
  364. mutate(Dimension = gsub('_diff', '', Dimension),
  365. Dimension = case_when(
  366. Dimension == "intero" ~ "Interoception",
  367. Dimension == "auditory" ~ "Hearing",
  368. Dimension == "olfactory" ~ "Olfaction",
  369. TRUE ~ str_to_title(Dimension)
  370. ))
  371. df_diffs_raw$Dimension = factor(df_diffs_raw$Dimension,
  372. levels = rev(c(
  373. 'Vision',
  374. 'Hearing',
  375. 'Olfaction',
  376. 'Taste',
  377. 'Interoception',
  378. 'Touch',
  379. 'Mouth_throat',
  380. 'Head',
  381. 'Torso',
  382. 'Hand_arm',
  383. 'Foot_leg'
  384. )))
  385. df_diffs_raw %>%
  386. filter(word == "market") %>%
  387. ggplot(aes(x = Dimension,
  388. y = Diff,
  389. fill = Dimension)) +
  390. geom_bar(stat = "summary") +
  391. geom_vline(xintercept = 0, linetype = "dotted") +
  392. theme_bw() +
  393. coord_flip() +
  394. labs(x = "Dimension",
  395. y = "Deviation from Lancaster Norms") +
  396. scale_fill_manual(values = viridisLite::viridis(11, option = "mako",
  397. begin = 0.8, end = 0.15)) +
  398. facet_wrap(~sentence) +
  399. theme(text = element_text(size=16)) +
  400. guides(fill = FALSE)
  401. df_paired_market <- df_dom_plus_sm %>%
  402. filter(word == "market") %>%
  403. select(sentence,
  404. Vision.M, Hearing.M, Olfaction.M, Taste.M, Interoception.M, Touch.M,
  405. Mouth_throat.M, Head.M, Torso.M, Hand_arm.M, Foot_leg.M,
  406. Visual.mean, Auditory.mean, Olfactory.mean, Gustatory.mean,
  407. Interoceptive.mean, Haptic.mean, Mouth.mean, Head.mean,
  408. Torso.mean, Hand_arm.mean, Foot_leg.mean) %>%
  409. pivot_longer(cols = -sentence, names_to = "raw_dim", values_to = "Rating") %>%
  410. mutate(
  411. Source = case_when(
  412. str_detect(raw_dim, "\\.M$") ~ "CS Norms (contextualized)",
  413. TRUE ~ "LS Norms (decontextualized)"
  414. ),
  415. Dimension = case_when(
  416. raw_dim %in% c("Vision.M", "Visual.mean") ~ "Vision",
  417. raw_dim %in% c("Hearing.M", "Auditory.mean") ~ "Hearing",
  418. raw_dim %in% c("Olfaction.M", "Olfactory.mean") ~ "Olfaction",
  419. raw_dim %in% c("Taste.M", "Gustatory.mean") ~ "Taste",
  420. raw_dim %in% c("Interoception.M", "Interoceptive.mean") ~ "Interoception",
  421. raw_dim %in% c("Touch.M", "Haptic.mean") ~ "Touch",
  422. raw_dim %in% c("Mouth_throat.M", "Mouth.mean") ~ "Mouth_throat",
  423. raw_dim %in% c("Head.M", "Head.mean") ~ "Head",
  424. raw_dim %in% c("Torso.M", "Torso.mean") ~ "Torso",
  425. raw_dim %in% c("Hand_arm.M", "Hand_arm.mean") ~ "Hand_arm",
  426. raw_dim %in% c("Foot_leg.M", "Foot_leg.mean") ~ "Foot_leg"
  427. )
  428. )
  429. df_paired_market$Dimension <- factor(df_paired_market$Dimension,
  430. levels = rev(c('Vision','Hearing','Olfaction','Taste','Interoception',
  431. 'Touch','Mouth_throat','Head','Torso','Hand_arm','Foot_leg')))
  432. df_paired_market %>%
  433. ggplot(aes(x = Dimension, y = Rating, fill = Source)) +
  434. geom_bar(stat = "summary", position = position_dodge(width = 0.8), width = 0.7) +
  435. theme_bw() +
  436. coord_flip() +
  437. labs(x = "Dimension", y = "Sensorimotor Strength", fill = NULL) +
  438. scale_fill_manual(values = c("CS Norms (contextualized)" = "#2c7a7b",
  439. "LS Norms (decontextualized)" = "#7c5e9c")) +
  440. facet_wrap(~sentence) +
  441. theme(text = element_text(size = 14), legend.position = "bottom")
  442. ```
  443. ### Distribution of deviations
  444. ```{r largest_dev}
  445. df_diffs_raw %>%
  446. mutate(abs_diff = abs(Diff)) %>%
  447. summarise(mean_abs_diff = mean(abs_diff),
  448. median_abs_diff = median(abs_diff),
  449. sd_abs_diff = sd(abs_diff),
  450. min_abs_diff = min(abs_diff),
  451. max_abs_diff = max(abs_diff))
  452. df_diffs_raw %>%
  453. mutate(abs_diff = abs(Diff)) %>%
  454. ggplot(aes(x = abs_diff)) +
  455. geom_histogram() +
  456. labs(x = "Absolute Deviation") +
  457. theme_minimal() +
  458. facet_wrap(~Dimension) +
  459. theme(text = element_text(size=16))
  460. df_by_sentence_raw = df_diffs_raw %>%
  461. mutate(abs_diff = abs(Diff)) %>%
  462. group_by(word, sentence) %>%
  463. summarise(mean_abs_diff = mean(abs_diff),
  464. min_diff = min(abs_diff),
  465. max_diff = max(abs_diff)) %>%
  466. ungroup()
  467. df_by_sentence_raw %>%
  468. summarise(mean_diff = mean(mean_abs_diff),
  469. median_abs_diff = median(mean_abs_diff),
  470. sd_abs_diff = sd(mean_abs_diff),
  471. min_abs_diff = min(mean_abs_diff),
  472. max_abs_diff = max(mean_abs_diff))
  473. df_by_sentence_raw %>%
  474. ggplot(aes(x = mean_abs_diff)) +
  475. geom_histogram() +
  476. labs(x = "Mean Absolute Deviation (By Sentence)") +
  477. theme_minimal() +
  478. theme(text = element_text(size=16))
  479. df_by_sentence_raw %>%
  480. ggplot(aes(x = max_diff)) +
  481. geom_histogram() +
  482. labs(x = "Max Absolute Deviation (By Sentence)") +
  483. theme_minimal() +
  484. theme(text = element_text(size=16))
  485. df_by_sentence_raw %>%
  486. arrange(desc(max_diff)) %>%
  487. head(10)
  488. df_by_word_raw = df_diffs_raw %>%
  489. mutate(abs_diff = abs(Diff)) %>%
  490. group_by(word) %>%
  491. summarise(mean_abs_diff = mean(abs_diff)) %>%
  492. ungroup()
  493. df_by_word_raw %>%
  494. summarise(mean_diff = mean(mean_abs_diff),
  495. median_abs_diff = median(mean_abs_diff),
  496. sd_abs_diff = sd(mean_abs_diff),
  497. min_abs_diff = min(mean_abs_diff),
  498. max_abs_diff = max(mean_abs_diff))
  499. df_diffs_raw %>%
  500. mutate(abs_diff = abs(Diff)) %>%
  501. group_by(word) %>%
  502. summarise(mean_abs_diff = mean(abs_diff),
  503. sd_abs_diff = sd(abs_diff),
  504. range_diff = max(abs_diff) - min(abs_diff)) %>%
  505. arrange(desc(mean_abs_diff)) %>%
  506. head(5)
  507. ```
  508. ### Individual words
  509. ```{r individual_words}
  510. df_diffs_raw = df_diffs_raw %>%
  511. mutate(sense = str_extract(context, "M[12]"), # "M1" or "M2"
  512. context_within_sense = str_extract(context, "[ab]")) # "a" or "b"
  513. df_diffs_raw %>%
  514. filter(word == "jam") %>%
  515. mutate(sentence = fct_reorder(sentence, as.numeric(factor(sense)))) %>%
  516. ggplot(aes(x = Dimension,
  517. y = Diff,
  518. fill = Dimension)) +
  519. geom_bar(stat = "summary") +
  520. geom_vline(xintercept = 0, linetype = "dotted") +
  521. theme_bw() +
  522. coord_flip() +
  523. labs(x = "Dimension",
  524. y = "Deviation from Lancaster Norms") +
  525. scale_fill_manual(values = viridisLite::viridis(11, option = "mako",
  526. begin = 0.8, end = 0.15)) +
  527. facet_wrap(~sentence) +
  528. theme(text = element_text(size=16)) +
  529. guides(fill = FALSE)
  530. df_diffs_raw %>%
  531. mutate(sentence = fct_reorder(sentence, as.numeric(factor(sense)))) %>%
  532. filter(word == "punch") %>%
  533. ggplot(aes(x = Dimension,
  534. y = Diff,
  535. fill = Dimension)) +
  536. geom_bar(stat = "summary") +
  537. geom_vline(xintercept = 0, linetype = "dotted") +
  538. theme_bw() +
  539. coord_flip() +
  540. labs(x = "Dimension",
  541. y = "Deviation from Lancaster Norms") +
  542. scale_fill_manual(values = viridisLite::viridis(11, option = "mako",
  543. begin = 0.8, end = 0.15)) +
  544. facet_wrap(~sentence) +
  545. theme(text = element_text(size=16)) +
  546. guides(fill = FALSE)
  547. df_diffs_raw %>%
  548. mutate(sentence = fct_reorder(sentence, as.numeric(factor(sense)))) %>%
  549. filter(word == "match") %>%
  550. ggplot(aes(x = Dimension,
  551. y = Diff,
  552. fill = Dimension)) +
  553. geom_bar(stat = "summary") +
  554. geom_vline(xintercept = 0, linetype = "dotted") +
  555. theme_bw() +
  556. coord_flip() +
  557. labs(x = "Dimension",
  558. y = "Deviation from Lancaster Norms") +
  559. scale_fill_manual(values = viridisLite::viridis(11, option = "mako",
  560. begin = 0.8, end = 0.15)) +
  561. facet_wrap(~sentence) +
  562. theme(text = element_text(size=16)) +
  563. guides(fill = FALSE)
  564. df_diffs_raw %>%
  565. mutate(sentence = fct_reorder(sentence, as.numeric(factor(sense)))) %>%
  566. filter(word == "pen") %>%
  567. ggplot(aes(x = Dimension,
  568. y = Diff,
  569. fill = Dimension)) +
  570. geom_bar(stat = "summary") +
  571. geom_vline(xintercept = 0, linetype = "dotted") +
  572. theme_bw() +
  573. coord_flip() +
  574. labs(x = "Dimension",
  575. y = "Deviation from Lancaster Norms") +
  576. scale_fill_manual(values = viridisLite::viridis(11, option = "mako",
  577. begin = 0.8, end = 0.15)) +
  578. facet_wrap(~sentence) +
  579. theme(text = element_text(size=16)) +
  580. guides(fill = FALSE)
  581. df_diffs_raw %>%
  582. mutate(sentence = fct_reorder(sentence, as.numeric(factor(sense)))) %>%
  583. filter(word == "band") %>%
  584. ggplot(aes(x = Dimension,
  585. y = Diff,
  586. fill = Dimension)) +
  587. geom_bar(stat = "summary") +
  588. geom_vline(xintercept = 0, linetype = "dotted") +
  589. theme_bw() +
  590. coord_flip() +
  591. labs(x = "Dimension",
  592. y = "Deviation from Lancaster Norms") +
  593. scale_fill_manual(values = viridisLite::viridis(11, option = "mako",
  594. begin = 0.8, end = 0.15)) +
  595. facet_wrap(~sentence) +
  596. theme(text = element_text(size=16)) +
  597. guides(fill = FALSE)
  598. ```
  599. ### Which dimensions have the largest deviations?
  600. ```{r}
  601. df_diffs_raw %>%
  602. mutate(abs_diff = abs(Diff)) %>%
  603. group_by(Dimension) %>%
  604. summarise(mean_abs_diff = mean(abs_diff),
  605. sd_abs_diff = sd(abs_diff),
  606. range_diff = max(abs_diff) - min(abs_diff)) %>%
  607. arrange(desc(range_diff))
  608. ```
  609. ## Z-scored deviations
  610. ```{r zscored_diffs}
  611. SD_FLOOR <- 0.1
  612. df_diffs_z_floored <- df_dom_plus_sm %>%
  613. mutate(
  614. vision_diff = (Vision.M - Visual.mean) / pmax(Visual.SD, SD_FLOOR),
  615. auditory_diff = (Hearing.M - Auditory.mean) / pmax(Auditory.SD, SD_FLOOR),
  616. intero_diff = (Interoception.M - Interoceptive.mean) / pmax(Interoceptive.SD, SD_FLOOR),
  617. olfactory_diff = (Olfaction.M - Olfactory.mean) / pmax(Olfactory.SD, SD_FLOOR),
  618. touch_diff = (Touch.M - Haptic.mean) / pmax(Haptic.SD, SD_FLOOR),
  619. taste_diff = (Taste.M - Gustatory.mean) / pmax(Gustatory.SD, SD_FLOOR),
  620. torso_diff = (Torso.M - Torso.mean) / pmax(Torso.SD.LS, SD_FLOOR),
  621. hand_arm_diff = (Hand_arm.M - Hand_arm.mean) / pmax(Hand_arm.SD.LS, SD_FLOOR),
  622. foot_leg_diff = (Foot_leg.M - Foot_leg.mean) / pmax(Foot_leg.SD.LS, SD_FLOOR),
  623. head_diff = (Head.M - Head.mean) / pmax(Head.SD.LS, SD_FLOOR),
  624. mouth_throat_diff = (Mouth_throat.M - Mouth.mean) / pmax(Mouth.SD, SD_FLOOR)
  625. ) %>%
  626. pivot_longer(cols = ends_with("_diff"),
  627. names_to = "Dimension", values_to = "Diff") %>%
  628. mutate(Dimension = gsub('_diff', '', Dimension),
  629. Dimension = case_when(
  630. Dimension == "intero" ~ "Interoception",
  631. Dimension == "auditory" ~ "Hearing",
  632. Dimension == "olfactory" ~ "Olfaction",
  633. TRUE ~ str_to_title(Dimension)
  634. )) %>%
  635. select(word, sentence, Dimension, Diff, context, dominance)
  636. # Distribution
  637. sentence_dev_floored <- df_diffs_z_floored %>%
  638. mutate(abs_diff = abs(Diff)) %>%
  639. group_by(word, sentence) %>%
  640. summarise(mean_abs_diff = mean(abs_diff, na.rm = TRUE),
  641. max_abs_diff = max(abs_diff),
  642. .groups = "drop")
  643. sentence_dev_floored %>%
  644. ggplot(aes(x = mean_abs_diff)) +
  645. geom_histogram(bins = 30) +
  646. geom_vline(xintercept = 1, linetype = "dashed", color = "red") +
  647. labs(x = "Mean Absolute Standardized Deviation (By Sentence)",
  648. y = "Count of sentences") +
  649. theme_minimal() +
  650. theme(text = element_text(size=16))
  651. sentence_dev_floored %>%
  652. ggplot(aes(x = max_abs_diff)) +
  653. geom_histogram(bins = 30) +
  654. geom_vline(xintercept = 1, linetype = "dashed", color = "red") +
  655. labs(x = "Max Absolute Standardized Deviation (By Sentence)",
  656. y = "Count of sentences") +
  657. theme_minimal() +
  658. theme(text = element_text(size=16))
  659. ### Overall % sentences more than one on average
  660. sentence_dev_floored %>%
  661. mutate(greater_than_one = mean_abs_diff > 1) %>%
  662. summarise(prop_more_than_one = mean(greater_than_one),
  663. max_dev = max(mean_abs_diff))
  664. ### Overall % sentences more than one maximally
  665. quantile(sentence_dev_floored$max_abs_diff,
  666. probs = c(0.25, 0.5, 0.75, 0.9, 0.95, 0.99),
  667. na.rm = TRUE)
  668. ### Overall deviations by dimension
  669. df_diffs_z_floored %>%
  670. mutate(abs_diff = abs(Diff)) %>%
  671. ggplot(aes(x = abs_diff)) +
  672. geom_histogram() +
  673. labs(x = "Standardized Absolute Deviation") +
  674. geom_vline(xintercept = 1, linetype = "dashed", color = "red") +
  675. theme_minimal() +
  676. facet_wrap(~Dimension) +
  677. theme(text = element_text(size=16))
  678. ### Overall % of contexts exceeding 1 SD
  679. df_diffs_z_floored %>%
  680. mutate(abs_diff = abs(Diff)) %>%
  681. mutate(greater_than_one = abs_diff > 1) %>%
  682. summarise(prop_more_than_one = mean(greater_than_one),
  683. max_dev = max(abs_diff))
  684. #### Additional summary statistics
  685. df_diffs_z_floored %>%
  686. mutate(abs_diff = abs(Diff)) %>%
  687. summarise(mean_abs_diff = mean(abs_diff),
  688. median_abs_diff = median(abs_diff),
  689. sd_abs_diff = sd(abs_diff),
  690. min_abs_diff = min(abs_diff),
  691. max_abs_diff = max(abs_diff))
  692. df_by_sentence_z_floored = df_diffs_z_floored %>%
  693. mutate(abs_diff = abs(Diff)) %>%
  694. group_by(word, sentence) %>%
  695. summarise(mean_abs_diff = mean(abs_diff)) %>%
  696. ungroup()
  697. df_by_sentence_z_floored %>%
  698. summarise(mean_diff = mean(mean_abs_diff),
  699. median_abs_diff = median(mean_abs_diff),
  700. sd_abs_diff = sd(mean_abs_diff),
  701. min_abs_diff = min(mean_abs_diff),
  702. max_abs_diff = max(mean_abs_diff))
  703. df_by_word_z_floored = df_diffs_z_floored %>%
  704. mutate(abs_diff = abs(Diff)) %>%
  705. group_by(word) %>%
  706. summarise(mean_abs_diff = mean(abs_diff)) %>%
  707. ungroup()
  708. df_by_word_z_floored %>%
  709. summarise(mean_diff = mean(mean_abs_diff),
  710. median_abs_diff = median(mean_abs_diff),
  711. sd_abs_diff = sd(mean_abs_diff),
  712. min_abs_diff = min(mean_abs_diff),
  713. max_abs_diff = max(mean_abs_diff))
  714. ```
  715. ### Compare to raw differences
  716. ```{r}
  717. # Raw (unstandardized) mean absolute deviation per sentence
  718. sentence_dev_raw <- df_diffs_raw %>%
  719. mutate(abs_diff = abs(Diff)) %>%
  720. group_by(word, sentence, context) %>%
  721. summarise(mean_abs_diff = mean(abs_diff, na.rm = TRUE),
  722. .groups = "drop") %>%
  723. mutate(version = "Raw")
  724. # Standardized (floored) mean absolute deviation per sentence
  725. sentence_dev_z <- df_diffs_z_floored %>%
  726. mutate(abs_diff = abs(Diff)) %>%
  727. group_by(word, sentence, context) %>%
  728. summarise(mean_abs_diff = mean(abs_diff, na.rm = TRUE),
  729. .groups = "drop") %>%
  730. mutate(version = "Standardized")
  731. # By sentence
  732. sentence_compare <- sentence_dev_raw %>%
  733. select(word, sentence, context, raw = mean_abs_diff) %>%
  734. inner_join(sentence_dev_z %>%
  735. select(word, sentence, context, standardized = mean_abs_diff),
  736. by = c("word", "sentence", "context"))
  737. cor.test(sentence_compare$raw, sentence_compare$standardized, method = "spearman")
  738. cor.test(sentence_compare$raw, sentence_compare$standardized, method = "pearson")
  739. sentence_compare %>%
  740. ggplot(aes(x = raw, y = standardized)) +
  741. geom_point(alpha = 0.3) +
  742. geom_smooth(method = "lm", se = FALSE) +
  743. labs(x = "Raw mean absolute deviation",
  744. y = "Standardized mean absolute deviation",
  745. title = "By-sentence deviations") +
  746. theme_minimal() +
  747. theme(text = element_text(size=16))
  748. ```
  749. ## Within-word range
  750. Here, we calculate the *range* of sensorimotor associatoins for each dimension across the contexts of use for each word.
  751. ### Across all contexts
  752. ```{r within_word_range}
  753. df_rawc_with_norms = read_csv("../../data/processed/sentence_pairs_with_sensorimotor_distance.csv") %>%
  754. drop_na(sensorimotor_distance) %>%
  755. select(sensorimotor_distance, action_distance, perceptual_distance,
  756. word, same, ambiguity_type, sentence1, sentence2, mean_relatedness, Class, version)
  757. nrow(df_rawc_with_norms)
  758. word_dim_range <- df_dom_plus_sm %>%
  759. pivot_longer(cols = c(Vision.M, Hearing.M, Olfaction.M, Taste.M,
  760. Interoception.M, Touch.M, Mouth_throat.M,
  761. Head.M, Hand_arm.M, Foot_leg.M, Torso.M),
  762. names_to = "Dimension", values_to = "CS_value") %>%
  763. mutate(Dimension = gsub("\\.M$", "", Dimension)) %>%
  764. group_by(word, Dimension) %>%
  765. summarise(within_range = max(CS_value, na.rm = TRUE) - min(CS_value, na.rm = TRUE),
  766. .groups = "drop")
  767. # Aggregate to word level (mean range across dimensions)
  768. word_range <- word_dim_range %>%
  769. group_by(word) %>%
  770. summarise(mean_range = mean(within_range, na.rm = TRUE),
  771. max_range = max(within_range, na.rm = TRUE),
  772. .groups = "drop") %>%
  773. arrange(desc(mean_range))
  774. print(head(word_range, 15))
  775. print(tail(word_range, 15))
  776. # By dimension
  777. word_dim_range %>%
  778. ggplot(aes(x = within_range)) +
  779. geom_histogram(bins = 30) +
  780. labs(x = "Within-Word Range of CS Ratings (across contexts)",
  781. y = "Count of words") +
  782. theme_minimal() +
  783. facet_wrap(~Dimension)
  784. # Distribution
  785. word_range %>%
  786. ggplot(aes(x = mean_range)) +
  787. geom_histogram(bins = 30) +
  788. labs(x = "Mean Within-Word Range of CS Ratings (across contexts)",
  789. y = "Count of words") +
  790. theme_minimal()
  791. # Summary stats
  792. word_range %>%
  793. summarise(mean = mean(mean_range),
  794. sd = sd(mean_range),
  795. median = median(mean_range),
  796. min = min(mean_range),
  797. max = max(mean_range))
  798. # Get per-word ambiguity type
  799. word_ambiguity <- df_rawc_with_norms %>%
  800. distinct(word, ambiguity_type)
  801. # Merge with within-word range
  802. word_range_with_amb <- word_range %>%
  803. inner_join(word_ambiguity, by = "word")
  804. # Summary stats by ambiguity type
  805. word_range_with_amb %>%
  806. group_by(ambiguity_type) %>%
  807. summarise(mean = mean(mean_range),
  808. sd = sd(mean_range),
  809. median = median(mean_range),
  810. n = n(),
  811. .groups = "drop")
  812. # Test the difference
  813. summary(lm(data = word_range_with_amb,
  814. mean_range ~ ambiguity_type))
  815. t.test(mean_range ~ ambiguity_type, data = word_range_with_amb)
  816. # Visualize
  817. word_range_with_amb %>%
  818. ggplot(aes(x = ambiguity_type, y = mean_range, fill = ambiguity_type)) +
  819. geom_boxplot(alpha = 0.6) +
  820. geom_jitter(width = 0.15, alpha = 0.3) +
  821. labs(x = "Ambiguity Type",
  822. y = "Mean Within-Word Range of CS Ratings") +
  823. theme_minimal() +
  824. theme(legend.position = "none")
  825. ```
  826. ### Between vs. Within sense
  827. ```{r within_word_range_senses}
  828. # Need sense info from context column (M1/M2)
  829. word_range_decomposed <- df_dom_plus_sm %>%
  830. mutate(sense = str_extract(context, "M[12]")) %>%
  831. pivot_longer(cols = c(Vision.M, Hearing.M, Olfaction.M, Taste.M,
  832. Interoception.M, Touch.M, Mouth_throat.M,
  833. Head.M, Hand_arm.M, Foot_leg.M, Torso.M),
  834. names_to = "Dimension", values_to = "CS_value") %>%
  835. mutate(Dimension = gsub("\\.M$", "", Dimension)) %>%
  836. group_by(word, Dimension, sense) %>%
  837. summarise(sense_mean = mean(CS_value),
  838. sense_within_range = max(CS_value) - min(CS_value),
  839. .groups = "drop") %>%
  840. group_by(word, Dimension) %>%
  841. summarise(between_sense_range = max(sense_mean) - min(sense_mean),
  842. within_sense_range = mean(sense_within_range),
  843. .groups = "drop") %>%
  844. group_by(word) %>%
  845. summarise(mean_between_sense = mean(between_sense_range),
  846. mean_within_sense = mean(within_sense_range),
  847. max_between_sense = max(between_sense_range),
  848. max_within_sense = max(within_sense_range),
  849. .groups = "drop")
  850. word_range_decomposed %>%
  851. inner_join(word_ambiguity, by = "word") %>%
  852. group_by(ambiguity_type) %>%
  853. summarise(between = mean(mean_between_sense),
  854. within = mean(mean_within_sense),,
  855. between_max = mean(max_between_sense),
  856. within_max = mean(max_within_sense),
  857. n = n())
  858. # Test each separately
  859. t.test(mean_between_sense ~ ambiguity_type,
  860. data = word_range_decomposed %>% inner_join(word_ambiguity, by = "word"))
  861. t.test(mean_within_sense ~ ambiguity_type,
  862. data = word_range_decomposed %>% inner_join(word_ambiguity, by = "word"))
  863. # Test max
  864. t.test(max_between_sense ~ ambiguity_type,
  865. data = word_range_decomposed %>% inner_join(word_ambiguity, by = "word"))
  866. t.test(max_within_sense ~ ambiguity_type,
  867. data = word_range_decomposed %>% inner_join(word_ambiguity, by = "word"))
  868. word_range_decomposed_long = word_range_decomposed %>%
  869. inner_join(word_ambiguity, by = "word") %>%
  870. pivot_longer(cols = c(mean_between_sense, mean_within_sense),
  871. names_to = "comparison",
  872. values_to = "word_range")
  873. summary(lmer(data = word_range_decomposed_long,
  874. word_range ~ comparison + (1|word)))
  875. summary(lmer(data = word_range_decomposed_long,
  876. word_range ~ ambiguity_type*comparison + (1|word)))
  877. word_range_decomposed_long %>%
  878. mutate(comparison = case_when(
  879. comparison == "mean_between_sense" ~ "Different Sense",
  880. comparison == "mean_within_sense" ~ "Same Sense"
  881. )) %>%
  882. ggplot(aes(x = word_range,
  883. y = ambiguity_type,
  884. fill = comparison)) +
  885. geom_density_ridges2(aes(height = ..density..),
  886. color = NA,
  887. alpha = 0.5,
  888. scale=0.85,
  889. stat="density") +
  890. labs(x = "Mean Within-Word Range of CS Ratings",
  891. y = NULL,
  892. fill = NULL) +
  893. theme_minimal() +
  894. scale_fill_manual(
  895. values = c("Same Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[1],
  896. "Different Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[2])
  897. ) +
  898. theme(text = element_text(size = 20))
  899. word_range_decomposed_long = word_range_decomposed %>%
  900. inner_join(word_ambiguity, by = "word") %>%
  901. pivot_longer(cols = c(max_between_sense, max_within_sense),
  902. names_to = "comparison",
  903. values_to = "word_range")
  904. summary(lmer(data = word_range_decomposed_long,
  905. word_range ~ comparison + (1|word)))
  906. summary(lmer(data = word_range_decomposed_long,
  907. word_range ~ ambiguity_type*comparison + (1|word)))
  908. word_range_decomposed_long %>%
  909. mutate(comparison = case_when(
  910. comparison == "max_between_sense" ~ "Different Sense",
  911. comparison == "max_within_sense" ~ "Same Sense"
  912. )) %>%
  913. ggplot(aes(x = word_range,
  914. y = ambiguity_type,
  915. fill = comparison)) +
  916. geom_density_ridges2(aes(height = ..density..),
  917. color = NA,
  918. alpha = 0.5,
  919. scale=0.85,
  920. stat="density") +
  921. labs(x = "Max Within-Word Range of CS Ratings",
  922. y = NULL,
  923. fill = NULL) +
  924. theme_minimal() +
  925. scale_fill_manual(
  926. values = c("Same Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[1],
  927. "Different Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[2])
  928. ) +
  929. theme(text = element_text(size = 20))
  930. ```
  931. # How does dominance relate to deviation from the LS Norms?
  932. This question can in turn be decomposed into two questions:
  933. First, are more dominant senses **closer** to the LS Norms overall? We might expect this to be the case if the LS Norms reflect the dominant sense; that is, when people rate the sensorimotor properties of a decontextualized word, they might be more likely to index properties associated with the most dominant contexts or meanings of that word.
  934. And the answer is **yes**: more dominant senses are indeed more *similar* (less distant) from the Lancaster norm in terms of their sensorimotor profile.
  935. ```{r dominance_ls_deviation}
  936. model_with_dominance = lmer(data = df_dom_plus_sm,
  937. distance_to_lancaster ~ dominance + (1 | word),
  938. REML = FALSE)
  939. model_no_dominance = lmer(data = df_dom_plus_sm,
  940. distance_to_lancaster ~ (1 | word),
  941. REML = FALSE)
  942. summary(model_with_dominance)
  943. anova(model_with_dominance, model_no_dominance)
  944. df_dom_plus_sm %>%
  945. ggplot(aes(x = dominance,
  946. y = distance_to_lancaster)) +
  947. geom_point(alpha = .5) +
  948. geom_smooth(method = "lm") +
  949. labs(x = "Dominance",
  950. y = "Cosine Distance to Decontextualized LS Norm") +
  951. theme_bw()
  952. cor.test(df_dom_plus_sm$dominance, df_dom_plus_sm$distance_to_lancaster)
  953. ```
  954. And second: does dominance predict the **direction** of difference?
  955. The earlier analysis of dominance suggests that more dominant senses are more concrete than less dominant senses. Thus, we might expect that more dominant senses are also more concrete on average than the decontextualized norms.
  956. We find that this is true: that is, more dominant senses are *more concrete* on average than the LS norm.
  957. ```{r}
  958. df_diffs_avg = df_diffs_raw %>%
  959. group_by(word, sentence, context) %>%
  960. summarise(mean_diff = mean(Diff))
  961. df_diffs_avg = df_diffs_avg %>%
  962. left_join(df_dom_plus_sm)
  963. model_with_dominance = lmer(data = df_diffs_avg,
  964. mean_diff ~ dominance + (1 | word),
  965. REML = FALSE)
  966. model_no_dominance = lmer(data = df_diffs_avg,
  967. mean_diff ~ (1 | word),
  968. REML = FALSE)
  969. summary(model_with_dominance)
  970. anova(model_with_dominance, model_no_dominance)
  971. cor.test(df_diffs_avg$dominance, df_diffs_avg$mean_diff)
  972. ```
  973. # Predicting relatedness
  974. Next, we ask about the **sensorimotor distance** between two sentence pairs, and whether it correlates both with `same/different sense` and the `mean_relatedness` judgments for those sentence pairs.
  975. Here, we load a version of the dataset that also contains a *baseline* measure: the sensorimotor distance as calculated using a bag-of-words approach (i.e., using the original Lancaster Norms).
  976. ## Load data
  977. ```{r}
  978. df_rawc_with_norms = read_csv("../../data/processed/sentence_pairs_with_sensorimotor_distance.csv") %>%
  979. drop_na(sensorimotor_distance) %>%
  980. select(sensorimotor_distance, action_distance, perceptual_distance,
  981. word, same, ambiguity_type, sentence1, sentence2, mean_relatedness, Class, version)
  982. nrow(df_rawc_with_norms)
  983. ```
  984. ## Load English RAW-C data
  985. ```{r}
  986. df_bert = read_csv("../../data/processed/models_english/rawc-distances_model-bert-base-uncased.csv") %>%
  987. mutate(Model = "BERT-base-uncased",
  988. Multilingual = "Monolingual")
  989. df_bert_cased = read_csv("../../data/processed/models_english/rawc-distances_model-bert-base-cased.csv") %>%
  990. mutate(Model = "BERT-base-cased",
  991. Multilingual = "Monolingual")
  992. df_xlm = read_csv("../../data/processed/models_english/rawc-distances_model-xlm-roberta-base.csv") %>%
  993. mutate(Model = "XLM-RoBERTa",
  994. Multilingual = "Multilingual")
  995. df_ab1 = read_csv("../../data/processed/models_english/rawc-distances_model-albert-base-v1.csv") %>%
  996. mutate(Model = "ALBERT-base-v1",
  997. Multilingual = "Monolingual")
  998. df_ab2 = read_csv("../../data/processed/models_english/rawc-distances_model-albert-base-v2.csv") %>%
  999. mutate(Model = "ALBERT-base-v2",
  1000. Multilingual = "Monolingual")
  1001. df_al = read_csv("../../data/processed/models_english/rawc-distances_model-albert-large-v2.csv") %>%
  1002. mutate(Model = "ALBERT-large-v2",
  1003. Multilingual = "Monolingual")
  1004. df_axl = read_csv("../../data/processed/models_english/rawc-distances_model-albert-xlarge-v2.csv") %>%
  1005. mutate(Model = "ALBERT-xlarge-v2",
  1006. Multilingual = "Monolingual")
  1007. df_axxl = read_csv("../../data/processed/models_english/rawc-distances_model-albert-xxlarge-v2.csv") %>%
  1008. mutate(Model = "ALBERT-xxlarge-v2",
  1009. Multilingual = "Monolingual")
  1010. df_rb = read_csv("../../data/processed/models_english/rawc-distances_model-roberta-base.csv") %>%
  1011. mutate(Model = "RoBERTa-base",
  1012. Multilingual = "Monolingual")
  1013. df_rl = read_csv("../../data/processed/models_english/rawc-distances_model-roberta-large.csv") %>%
  1014. mutate(Model = "RoBERTa-large",
  1015. Multilingual = "Monolingual")
  1016. df_db = read_csv("../../data/processed/models_english/rawc-distances_model-distilbert-base-uncased.csv") %>%
  1017. mutate(Model = "DistilBERT",
  1018. Multilingual = "Monolingual")
  1019. df_mb = read_csv("../../data/processed/models_english/rawc-distances_model-bert-base-multilingual-cased.csv") %>%
  1020. mutate(Model = "Multilingual BERT",
  1021. Multilingual = "Multilingual")
  1022. df_all = df_bert %>%
  1023. bind_rows(df_bert_cased) %>%
  1024. bind_rows(df_xlm) %>%
  1025. bind_rows(df_ab1) %>%
  1026. bind_rows(df_ab2) %>%
  1027. bind_rows(df_al) %>%
  1028. bind_rows(df_axl) %>%
  1029. bind_rows(df_axxl) %>%
  1030. bind_rows(df_rb) %>%
  1031. bind_rows(df_rl) %>%
  1032. bind_rows(df_db) %>%
  1033. bind_rows(df_mb)
  1034. df_merged = df_rawc_with_norms %>%
  1035. inner_join(df_all)
  1036. ```
  1037. ## Predicting same/different sense
  1038. ### Sensorimotor distance and sense boundaries
  1039. ```{r sm_sense}
  1040. df_rawc_with_norms = df_rawc_with_norms %>%
  1041. mutate(Same = case_when(
  1042. same == TRUE ~ "Same Sense",
  1043. same == FALSE ~ "Different Sense"
  1044. ))
  1045. df_rawc_with_norms %>%
  1046. ggplot(aes(x = sensorimotor_distance,
  1047. y = ambiguity_type,
  1048. fill = Same)) +
  1049. geom_density_ridges2(aes(height = ..density..),
  1050. color = NA,
  1051. alpha = 0.5,
  1052. scale=0.85,
  1053. stat="density") +
  1054. labs(x = "Sensorimotor Distance",
  1055. y = NULL,
  1056. fill = NULL) +
  1057. scale_fill_manual(
  1058. values = c("Same Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[1],
  1059. "Different Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[2])
  1060. ) +
  1061. theme_minimal() +
  1062. theme(text = element_text(size = 20))
  1063. model_full = lmer(data = df_rawc_with_norms,
  1064. sensorimotor_distance ~ same +
  1065. (1 + same | word),
  1066. control=lmerControl(optimizer="bobyqa"),
  1067. REML = FALSE)
  1068. model_reduced = lmer(data = df_rawc_with_norms,
  1069. sensorimotor_distance ~ # same +
  1070. (1 + same | word),
  1071. control=lmerControl(optimizer="bobyqa"),
  1072. REML = FALSE)
  1073. summary(model_full)
  1074. anova(model_full, model_reduced)
  1075. model_full_at = lmer(data = df_rawc_with_norms,
  1076. sensorimotor_distance ~ same * ambiguity_type +
  1077. (1 + same | word),
  1078. control=lmerControl(optimizer="bobyqa"),
  1079. REML = FALSE)
  1080. model_reduced_at = lmer(data = df_rawc_with_norms,
  1081. sensorimotor_distance ~ same + ambiguity_type +
  1082. (1 + same | word),
  1083. control=lmerControl(optimizer="bobyqa"),
  1084. REML = FALSE)
  1085. summary(model_full_at)
  1086. anova(model_full_at, model_reduced_at)
  1087. df_rawc_with_norms %>%
  1088. group_by(same, ambiguity_type) %>%
  1089. summarise(mean = mean(sensorimotor_distance),
  1090. sd = sd(sensorimotor_distance))
  1091. ```
  1092. ### Predicting sense boundaries
  1093. Here, we run the actual analysis:
  1094. ```{r aic_sense_boundaries}
  1095. model_sm_only <- glm(same ~ sensorimotor_distance,
  1096. data = df_rawc_with_norms,
  1097. family = binomial)
  1098. aic_sm_baseline <- AIC(model_sm_only)
  1099. print(paste("AIC:", round(aic_sm_baseline, 1)))
  1100. ### Get aic
  1101. aic_results <- df_merged %>%
  1102. group_by(n_params, Model, Layer) %>%
  1103. summarize(
  1104. # BERT only
  1105. aic_bert = {
  1106. model <- glm(Same_sense ~ Distance, family = binomial)
  1107. AIC(model)
  1108. },
  1109. # Combined
  1110. aic_combined = {
  1111. model <- glm(Same_sense ~ Distance + sensorimotor_distance, family = binomial)
  1112. AIC(model)
  1113. },
  1114. aic_sm = {
  1115. model <- glm(Same_sense ~ sensorimotor_distance, family = binomial)
  1116. AIC(model)
  1117. },
  1118. .groups = "drop"
  1119. ) %>%
  1120. mutate(
  1121. # AIC Delta
  1122. aic_bert_vs_sm = aic_bert - aic_sm_baseline,
  1123. aic_combined_vs_sm = aic_sm_baseline - aic_combined,
  1124. aic_combined_vs_dist = aic_bert - aic_combined
  1125. )
  1126. # ============================================================
  1127. # 3. Best layer per model
  1128. # ============================================================
  1129. best_layers <- aic_results %>%
  1130. group_by(Model, n_params) %>%
  1131. slice_min(aic_bert, n = 1) %>%
  1132. select(Model, Layer,
  1133. aic_bert_vs_sm, aic_combined_vs_sm, aic_combined_vs_dist,
  1134. aic_sm, aic_bert, aic_combined)
  1135. best_layers = best_layers %>%
  1136. mutate(aic_bert_vs_sm2 = aic_bert - aic_sm_baseline)
  1137. print(best_layers)
  1138. ## AIC difference from SM baseline
  1139. best_layers %>%
  1140. select(Model, aic_bert_vs_sm, aic_combined_vs_sm) %>%
  1141. pivot_longer(cols = c(aic_bert_vs_sm, aic_combined_vs_sm),
  1142. names_to = "type", values_to = "aic_diff") %>%
  1143. mutate(type = ifelse(type == "aic_bert_vs_sm",
  1144. "Distributional only",
  1145. "Distributional + Sensorimotor"),
  1146. type = factor(type, levels = c("Distributional only",
  1147. "Distributional + Sensorimotor")),
  1148. Model = fct_reorder(Model, aic_diff)) %>%
  1149. ggplot(aes(x = aic_diff, y = Model, fill = type)) +
  1150. geom_col(position = "dodge") +
  1151. geom_vline(xintercept = 0, linetype = "dashed", color = "red") +
  1152. labs(x = "ΔAIC vs. Sensorimotor baseline",
  1153. y = NULL,
  1154. fill = NULL) +
  1155. scale_fill_manual(values = c("Distributional only" = "gray60",
  1156. "Distributional + Sensorimotor" = "steelblue")) +
  1157. theme_minimal() +
  1158. theme(text = element_text(size = 16),
  1159. legend.position = "bottom")
  1160. best_layers %>%
  1161. ungroup() %>%
  1162. select(Model, aic_bert, aic_combined) %>%
  1163. mutate(across(starts_with("aic"), round)) %>%
  1164. kbl(col.names = c("Model", "Distributional", "Hybrid"),
  1165. format = "latex",
  1166. booktabs = TRUE,
  1167. caption = "AIC values for predicting sense boundaries using Distributional Distance alone or both Distributional Distance and Sensorimotor Distance (AIC for Sensorimotor Distance was 707)",
  1168. label = "aic") %>%
  1169. kable_styling(latex_options = "hold_position")
  1170. aic_results %>%
  1171. mutate(dist_better_than_sm = aic_bert_vs_sm > 4,
  1172. hybrid_better_than_sm = aic_combined_vs_sm > 4,
  1173. hybrid_better_than_dist = aic_combined_vs_dist > 4) %>%
  1174. summarise(mean(dist_better_than_sm),
  1175. mean(hybrid_better_than_sm),
  1176. mean(hybrid_better_than_dist))
  1177. best_layers %>%
  1178. ungroup() %>%
  1179. mutate(dist_better_than_sm = aic_bert_vs_sm < -4,
  1180. hybrid_better_than_sm = aic_combined_vs_sm > 4,
  1181. hybrid_better_than_dist = aic_combined_vs_dist > 4) %>%
  1182. summarise(mean(dist_better_than_sm),
  1183. mean(hybrid_better_than_sm),
  1184. mean(hybrid_better_than_dist))
  1185. ### Visualize raw AIC
  1186. aic_summary_full <- best_layers %>%
  1187. select(Model, Layer, aic_sm, aic_bert, aic_combined) %>%
  1188. pivot_longer(cols = c(aic_sm, aic_bert, aic_combined),
  1189. names_to = "type", values_to = "aic") %>%
  1190. mutate(type = case_when(
  1191. type == "aic_sm" ~ "Sensorimotor",
  1192. type == "aic_bert" ~ "Distributional",
  1193. type == "aic_combined" ~ "Hybrid"
  1194. ),
  1195. type = factor(type, levels = c("Sensorimotor", "Distributional", "Hybrid")))
  1196. # Get sensorimotor baseline (should be the same for all models)
  1197. aic_sm_baseline <- aic_summary_full %>%
  1198. filter(type == "Sensorimotor") %>%
  1199. pull(aic) %>%
  1200. unique()
  1201. # Filter to just distributional and hybrid
  1202. aic_summary_no_sm <- aic_summary_full %>%
  1203. filter(type != "Sensorimotor")
  1204. aic_summary_se_no_sm <- aic_summary_no_sm %>%
  1205. group_by(type) %>%
  1206. summarize(
  1207. mean_aic = mean(aic),
  1208. se_aic = sd(aic) / sqrt(n())
  1209. )
  1210. ggplot() +
  1211. # Sensorimotor baseline as dashed line
  1212. geom_hline(yintercept = aic_sm_baseline,
  1213. linetype = "dashed", color = "coral", linewidth = 1) +
  1214. # Raw points for each model/layer
  1215. geom_jitter(data = aic_summary_no_sm,
  1216. aes(x = type, y = aic, color = type),
  1217. alpha = 0.3, width = 0.1, size = 2) +
  1218. # Mean + SE
  1219. geom_point(data = aic_summary_se_no_sm,
  1220. aes(x = type, y = mean_aic, color = type),
  1221. size = 5) +
  1222. geom_errorbar(data = aic_summary_se_no_sm,
  1223. aes(x = type, ymin = mean_aic - se_aic,
  1224. ymax = mean_aic + se_aic, color = type),
  1225. width = 0.15, linewidth = 1.2) +
  1226. annotate("text", x = 1.5, y = aic_sm_baseline + 30,
  1227. label = "Sensorimotor", color = "coral", size = 5)+
  1228. labs(x = NULL,
  1229. y = "AIC (lower is better)",
  1230. title = "Predicting Sense Boundary") +
  1231. scale_color_manual(values = c("Distributional" = "gray60",
  1232. "Hybrid" = "steelblue")) +
  1233. theme_minimal() +
  1234. theme(text = element_text(size = 16),
  1235. legend.position = "none")
  1236. ```
  1237. ## Predicting relatedness
  1238. ### Comparing to each distributional relatedness measure
  1239. ```{r relatedness_comparison}
  1240. df_rawc_with_norms %>%
  1241. ggplot(aes(x = sensorimotor_distance,
  1242. y = mean_relatedness,
  1243. color = Same)) +
  1244. geom_point(alpha = .5) +
  1245. scale_color_manual(
  1246. values = c("Same Sense" = viridisLite::viridis(2, option = "mako",
  1247. begin = 0.8, end = 0.15)[1],
  1248. "Different Sense" = viridisLite::viridis(2, option = "mako",
  1249. begin = 0.8, end = 0.15)[2])
  1250. ) +
  1251. theme_minimal() +
  1252. labs(x = "Sensorimotor Distance",
  1253. y = "Mean Relatedness",
  1254. color = "") +
  1255. theme(text = element_text(size = 15),
  1256. legend.position="bottom")
  1257. cor(df_rawc_with_norms$sensorimotor_distance,
  1258. df_rawc_with_norms$mean_relatedness)
  1259. ```
  1260. ### Comparing AIC
  1261. ```{r aic_relatedness}
  1262. # ============================================================
  1263. # 2. For each model/layer: compute R² and AIC, difference from SM baseline
  1264. # ============================================================
  1265. relatedness_results <- df_merged %>%
  1266. group_by(Model, Layer, n_params, Multilingual) %>%
  1267. summarize(
  1268. # Sensorimotor only
  1269. r2_sm = {
  1270. model <- lm(mean_relatedness ~ sensorimotor_distance)
  1271. summary(model)$r.squared
  1272. },
  1273. aic_sm = {
  1274. model <- lm(mean_relatedness ~ sensorimotor_distance)
  1275. AIC(model)
  1276. },
  1277. # BERT only
  1278. r2_bert = {
  1279. model <- lm(mean_relatedness ~ Distance)
  1280. summary(model)$r.squared
  1281. },
  1282. aic_bert = {
  1283. model <- lm(mean_relatedness ~ Distance)
  1284. AIC(model)
  1285. },
  1286. # Combined
  1287. r2_combined = {
  1288. model <- lm(mean_relatedness ~ Distance + sensorimotor_distance)
  1289. summary(model)$r.squared
  1290. },
  1291. aic_combined = {
  1292. model <- lm(mean_relatedness ~ Distance + sensorimotor_distance)
  1293. AIC(model)
  1294. },
  1295. .groups = "drop"
  1296. ) %>%
  1297. mutate(
  1298. aic_bert_vs_sm = aic_bert - aic_sm,
  1299. aic_combined_vs_sm = aic_sm - aic_combined,
  1300. aic_combined_vs_dist = aic_bert - aic_combined
  1301. )
  1302. # ============================================================
  1303. # 3. Best layer per model
  1304. # ============================================================
  1305. best_layers <- relatedness_results %>%
  1306. group_by(Model, n_params, Multilingual) %>%
  1307. slice_max(r2_bert, n = 1) %>%
  1308. select(Model, Layer, n_params, Multilingual,
  1309. # r2_sm, r2_bert, r2_combined,
  1310. aic_bert_vs_sm, aic_sm, aic_bert, aic_combined,
  1311. aic_combined_vs_sm, aic_combined_vs_dist)
  1312. print(best_layers)
  1313. best_layers %>%
  1314. ungroup() %>%
  1315. select(Model, aic_bert, aic_combined) %>%
  1316. mutate(across(starts_with("aic"), round)) %>%
  1317. kbl(col.names = c("Model", "Distributional", "Hybrid"),
  1318. format = "latex",
  1319. booktabs = TRUE,
  1320. caption = "AIC values for predicting relatedness judgments using Distributional Distance alone or both Distributional Distance and Sensorimotor Distance (AIC for Sensorimotor Distance was 2134.78.)",
  1321. label = "aic") %>%
  1322. kable_styling(latex_options = "hold_position")
  1323. best_layers %>%
  1324. select(Model, aic_bert_vs_sm, aic_combined_vs_sm) %>%
  1325. pivot_longer(cols = c(aic_bert_vs_sm, aic_combined_vs_sm),
  1326. names_to = "type", values_to = "aic_diff") %>%
  1327. mutate(type = ifelse(type == "aic_bert_vs_sm", "Distributional only", "Distributional + Sensorimotor"),
  1328. type = factor(type, levels = c("Distributional only", "Distributional + Sensorimotor")),
  1329. Model = fct_reorder(Model, aic_diff)) %>%
  1330. ggplot(aes(x = aic_diff, y = Model, fill = type)) +
  1331. geom_col(position = "dodge") +
  1332. geom_vline(xintercept = 0, linetype = "dashed", color = "red") +
  1333. geom_vline(xintercept = 4, linetype = "dotted", color = "gray50") +
  1334. geom_vline(xintercept = -4, linetype = "dotted", color = "gray50") +
  1335. labs(x = "ΔAIC vs. Sensorimotor baseline",
  1336. y = NULL,
  1337. fill = NULL) +
  1338. scale_fill_manual(values = c("Distributional only" = "gray60",
  1339. "Distributional + Sensorimotor" = "steelblue")) +
  1340. theme_minimal() +
  1341. theme(text = element_text(size = 16),
  1342. legend.position = "bottom")
  1343. # ============================================================
  1344. # 7. Summary
  1345. # ============================================================
  1346. best_layers %>%
  1347. ungroup() %>%
  1348. mutate(dist_better_than_sm = aic_bert_vs_sm < -4,
  1349. hybrid_better_than_sm = aic_combined_vs_sm > 4,
  1350. hybrid_better_than_dist = aic_combined_vs_dist > 4) %>%
  1351. summarise(mean(dist_better_than_sm),
  1352. mean(hybrid_better_than_sm),
  1353. mean(hybrid_better_than_dist))
  1354. ### Visualize raw AIC
  1355. aic_summary_full <- best_layers %>%
  1356. select(Model, Layer, aic_sm, aic_bert, aic_combined) %>%
  1357. pivot_longer(cols = c(aic_sm, aic_bert, aic_combined),
  1358. names_to = "type", values_to = "aic") %>%
  1359. mutate(type = case_when(
  1360. type == "aic_sm" ~ "Sensorimotor",
  1361. type == "aic_bert" ~ "Distributional",
  1362. type == "aic_combined" ~ "Hybrid"
  1363. ),
  1364. type = factor(type, levels = c("Sensorimotor", "Distributional", "Hybrid")))
  1365. # Get sensorimotor baseline (should be the same for all models)
  1366. aic_sm_baseline <- aic_summary_full %>%
  1367. filter(type == "Sensorimotor") %>%
  1368. pull(aic) %>%
  1369. unique()
  1370. # Filter to just distributional and hybrid
  1371. aic_summary_no_sm <- aic_summary_full %>%
  1372. filter(type != "Sensorimotor")
  1373. aic_summary_se_no_sm <- aic_summary_no_sm %>%
  1374. group_by(type) %>%
  1375. summarize(
  1376. mean_aic = mean(aic),
  1377. se_aic = sd(aic) / sqrt(n())
  1378. )
  1379. ggplot() +
  1380. # Sensorimotor baseline as dashed line
  1381. geom_hline(yintercept = aic_sm_baseline,
  1382. linetype = "dashed", color = "coral", linewidth = 1) +
  1383. # Raw points for each model/layer
  1384. geom_jitter(data = aic_summary_no_sm,
  1385. aes(x = type, y = aic, color = type),
  1386. alpha = 0.3, width = 0.1, size = 2) +
  1387. # Mean + SE
  1388. geom_point(data = aic_summary_se_no_sm,
  1389. aes(x = type, y = mean_aic, color = type),
  1390. size = 5) +
  1391. geom_errorbar(data = aic_summary_se_no_sm,
  1392. aes(x = type, ymin = mean_aic - se_aic,
  1393. ymax = mean_aic + se_aic, color = type),
  1394. width = 0.15, linewidth = 1.2) +
  1395. annotate("text", x = 1.5, y = aic_sm_baseline + 30,
  1396. label = "Sensorimotor", color = "coral", size = 5)+
  1397. labs(x = NULL,
  1398. y = "AIC (lower is better)",
  1399. title = "Predicting Relatedness") +
  1400. scale_color_manual(values = c("Distributional" = "gray60",
  1401. "Hybrid" = "steelblue")) +
  1402. theme_minimal() +
  1403. theme(text = element_text(size = 16),
  1404. legend.position = "none")
  1405. ```
  1406. ## Correlation between sensorimotor and distributional distance
  1407. ```{r corr_dist_sm}
  1408. # Get best layer
  1409. df_by_layer = df_merged %>%
  1410. group_by(Model, Multilingual, Layer, n_params) %>%
  1411. summarise(r = cor(mean_relatedness, Distance, method = "pearson"),
  1412. r2 = r ** 2,
  1413. rho = cor(mean_relatedness, Distance, method = "spearman"),
  1414. count = n())
  1415. df_best_layer <- df_by_layer %>%
  1416. group_by(Model) %>%
  1417. slice_max(r2, n = 1) %>%
  1418. select(Model, Layer, r2)
  1419. # Filter for only best layers
  1420. df_best <- df_merged %>%
  1421. semi_join(df_best_layer, by = c("Model", "Layer"))
  1422. # Pivot wider
  1423. df_wide <- df_best %>%
  1424. mutate(`Sensorimotor Distance` = sensorimotor_distance) %>%
  1425. select(word, sentence1, sentence2, `Sensorimotor Distance`, Model, Distance) %>%
  1426. pivot_wider(names_from = Model, values_from = Distance)
  1427. # Compute correlation matrix
  1428. cor_matrix <- df_wide %>%
  1429. select(`Sensorimotor Distance`, where(is.numeric)) %>%
  1430. cor(use = "pairwise.complete.obs")
  1431. # Plot the correlation matrix
  1432. ggcorrplot(cor_matrix,
  1433. hc.order = FALSE,
  1434. method = "square") +
  1435. theme(
  1436. axis.text.x = element_text(size = 10, angle = 45, hjust = 1),
  1437. axis.text.y = element_text(size = 10)
  1438. )
  1439. ```
  1440. # Baseline with LS Norms
  1441. ```{r}
  1442. df_with_baseline = read_csv("../../data/processed/sentence_pairs_with_baseline.csv") %>%
  1443. drop_na(sensorimotor_distance) %>%
  1444. drop_na(baseline_distance)
  1445. nrow(df_with_baseline)
  1446. ```
  1447. ## Does sensorimotor distance predict relatedness above the baseline?
  1448. We also ask whether whether our measure of contextualized sensorimotor distance predicts relatedness above and beyond a baseline that simply considers the decontextualized Lancaster Sensorimotor Norms for the dismabiguating words in a sentence. (We find that it does.)
  1449. ```{r}
  1450. model_bow_sm = lmer(data = df_with_baseline,
  1451. mean_relatedness ~ baseline_distance + sensorimotor_distance +
  1452. (1| word),
  1453. REML = FALSE)
  1454. model_just_bow = lmer(data = df_with_baseline,
  1455. mean_relatedness ~ baseline_distance +
  1456. (1| word),
  1457. REML = FALSE)
  1458. model_just_sm = lmer(data = df_with_baseline,
  1459. mean_relatedness ~ sensorimotor_distance +
  1460. (1| word),
  1461. REML = FALSE)
  1462. summary(model_bow_sm)
  1463. anova(model_bow_sm, model_just_bow)
  1464. anova(model_bow_sm, model_just_sm)
  1465. ```
  1466. # Predicting Trott & Bergen (2023)
  1467. Here, we also ask whether sensorimotor distance can account for variance in RT and accuracy on a primed sensibility judgment task.
  1468. ```{r}
  1469. # ============================================================
  1470. # Load T&B 2023 data and merge with norms
  1471. # ============================================================
  1472. df_s1 <- read_csv("../../data/tb2023/polysemy_s1_final.csv")
  1473. df_s2 <- read_csv("../../data/tb2023/polysemy_s2_final.csv")
  1474. df_tb2023 <- bind_rows(df_s1, df_s2)
  1475. nrow(df_tb2023)
  1476. # Item-level aggregates
  1477. df_tb2023_agg_acc <- df_tb2023 %>%
  1478. group_by(word, same, version) %>%
  1479. summarise(accuracy = mean(correct_response), .groups = "drop")
  1480. df_tb2023_agg_rt <- df_tb2023 %>%
  1481. filter(correct_response == TRUE) %>%
  1482. group_by(word, same, version) %>%
  1483. summarise(mean_rt = mean(rt), .groups = "drop") %>%
  1484. mutate(mean_log_rt = log10(mean_rt))
  1485. # Merge with sensorimotor norms (for Part 1)
  1486. df_sm_acc <- df_rawc_with_norms %>%
  1487. select(word, same, version, sensorimotor_distance,
  1488. action_distance, perceptual_distance) %>%
  1489. inner_join(df_tb2023_agg_acc, by = c("word", "same", "version"))
  1490. df_sm_rt <- df_rawc_with_norms %>%
  1491. select(word, same, version, sensorimotor_distance,
  1492. action_distance, perceptual_distance) %>%
  1493. inner_join(df_tb2023_agg_rt, by = c("word", "same", "version"))
  1494. # Merge with full distributional sweep (for Part 2)
  1495. df_full_acc <- df_merged %>%
  1496. inner_join(df_tb2023_agg_acc, by = c("word", "same", "version"))
  1497. df_full_rt <- df_merged %>%
  1498. inner_join(df_tb2023_agg_rt, by = c("word", "same", "version"))
  1499. # Sanity check
  1500. nrow(df_sm_acc)
  1501. nrow(df_sm_rt)
  1502. nrow(df_full_acc)
  1503. nrow(df_full_rt)
  1504. ```
  1505. ## Part 1: SM on its own
  1506. ```{r}
  1507. mod_sm_acc = lmer(data = df_sm_acc,
  1508. accuracy ~ sensorimotor_distance +
  1509. (1 | word))
  1510. summary(mod_sm_acc)
  1511. mod_sm_rt = lmer(data = df_sm_rt,
  1512. mean_log_rt ~ sensorimotor_distance +
  1513. (1 | word))
  1514. summary(mod_sm_rt)
  1515. ```
  1516. ## Part 2: Comopare to distributional distance
  1517. Here, we compare sensorimotor distance to the best-performing distributional distance measure from the RAW-C analysis.
  1518. ```{r sm_dd_tb2023}
  1519. # --- Compute AIC for each model/layer: SM, Distributional, Hybrid ---
  1520. tb_acc_results <- df_full_acc %>%
  1521. group_by(Model, Layer, n_params, Multilingual) %>%
  1522. summarize(
  1523. r2_sm = summary(lm(accuracy ~ sensorimotor_distance))$r.squared,
  1524. aic_sm = AIC(lm(accuracy ~ sensorimotor_distance)),
  1525. r2_bert = summary(lm(accuracy ~ Distance))$r.squared,
  1526. aic_bert = AIC(lm(accuracy ~ Distance)),
  1527. r2_combined = summary(lm(accuracy ~ Distance + sensorimotor_distance))$r.squared,
  1528. aic_combined = AIC(lm(accuracy ~ Distance + sensorimotor_distance)),
  1529. .groups = "drop"
  1530. ) %>%
  1531. mutate(
  1532. aic_bert_vs_sm = aic_bert - aic_sm,
  1533. aic_combined_vs_sm = aic_sm - aic_combined,
  1534. aic_combined_vs_dist = aic_bert - aic_combined
  1535. )
  1536. tb_rt_results <- df_full_rt %>%
  1537. group_by(Model, Layer, n_params, Multilingual) %>%
  1538. summarize(
  1539. r2_sm = summary(lm(mean_log_rt ~ sensorimotor_distance))$r.squared,
  1540. aic_sm = AIC(lm(mean_log_rt ~ sensorimotor_distance)),
  1541. r2_bert = summary(lm(mean_log_rt ~ Distance))$r.squared,
  1542. aic_bert = AIC(lm(mean_log_rt ~ Distance)),
  1543. r2_combined = summary(lm(mean_log_rt ~ Distance + sensorimotor_distance))$r.squared,
  1544. aic_combined = AIC(lm(mean_log_rt ~ Distance + sensorimotor_distance)),
  1545. .groups = "drop"
  1546. ) %>%
  1547. mutate(
  1548. aic_bert_vs_sm = aic_bert - aic_sm,
  1549. aic_combined_vs_sm = aic_sm - aic_combined,
  1550. aic_combined_vs_dist = aic_bert - aic_combined
  1551. )
  1552. # --- Best layer per model (selected by best distributional R²) ---
  1553. best_layers_acc <- tb_acc_results %>%
  1554. group_by(Model, n_params, Multilingual) %>%
  1555. slice_max(r2_bert, n = 1) %>%
  1556. ungroup()
  1557. best_layers_rt <- tb_rt_results %>%
  1558. group_by(Model, n_params, Multilingual) %>%
  1559. slice_max(r2_bert, n = 1) %>%
  1560. ungroup()
  1561. print(best_layers_acc)
  1562. print(best_layers_rt)
  1563. # --- Summary counts ---
  1564. best_layers_acc %>%
  1565. mutate(dist_better_than_sm = aic_bert_vs_sm < -4,
  1566. hybrid_better_than_sm = aic_combined_vs_sm > 4,
  1567. hybrid_better_than_dist = aic_combined_vs_dist > 4) %>%
  1568. summarise(across(c(dist_better_than_sm,
  1569. hybrid_better_than_sm,
  1570. hybrid_better_than_dist), mean))
  1571. best_layers_rt %>%
  1572. mutate(dist_better_than_sm = aic_bert_vs_sm < -4,
  1573. hybrid_better_than_sm = aic_combined_vs_sm > 4,
  1574. hybrid_better_than_dist = aic_combined_vs_dist > 4) %>%
  1575. summarise(across(c(dist_better_than_sm,
  1576. hybrid_better_than_sm,
  1577. hybrid_better_than_dist), mean))
  1578. # --- LaTeX tables ---
  1579. best_layers_acc %>%
  1580. select(Model, aic_bert, aic_combined) %>%
  1581. mutate(across(starts_with("aic"), round)) %>%
  1582. kbl(col.names = c("Model", "Distributional", "Hybrid"),
  1583. format = "latex", booktabs = TRUE,
  1584. caption = sprintf("AIC values for predicting accuracy (Trott \\& Bergen, 2023). AIC for Sensorimotor Distance alone was %.2f.",
  1585. unique(best_layers_acc$aic_sm)),
  1586. label = "aic_tb_acc") %>%
  1587. kable_styling(latex_options = "hold_position")
  1588. best_layers_rt %>%
  1589. select(Model, aic_bert, aic_combined) %>%
  1590. mutate(across(starts_with("aic"), round)) %>%
  1591. kbl(col.names = c("Model", "Distributional", "Hybrid"),
  1592. format = "latex", booktabs = TRUE,
  1593. caption = sprintf("AIC values for predicting log RT (Trott \\& Bergen, 2023). AIC for Sensorimotor Distance alone was %.2f.",
  1594. unique(best_layers_rt$aic_sm)),
  1595. label = "aic_tb_rt") %>%
  1596. kable_styling(latex_options = "hold_position")
  1597. # ============================================================
  1598. # Figures: AIC dot plots (matching your existing style)
  1599. # ============================================================
  1600. # --- Accuracy ---
  1601. aic_long_acc <- best_layers_acc %>%
  1602. select(Model, Layer, aic_sm, aic_bert, aic_combined) %>%
  1603. pivot_longer(cols = c(aic_sm, aic_bert, aic_combined),
  1604. names_to = "type", values_to = "aic") %>%
  1605. mutate(type = case_when(
  1606. type == "aic_sm" ~ "Sensorimotor",
  1607. type == "aic_bert" ~ "Distributional",
  1608. type == "aic_combined" ~ "Hybrid"
  1609. ),
  1610. type = factor(type, levels = c("Sensorimotor", "Distributional", "Hybrid")))
  1611. baseline_acc <- aic_long_acc %>% filter(type == "Sensorimotor") %>% pull(aic) %>% unique()
  1612. aic_long_acc_no_sm <- aic_long_acc %>% filter(type != "Sensorimotor")
  1613. aic_se_acc <- aic_long_acc_no_sm %>%
  1614. group_by(type) %>%
  1615. summarize(mean_aic = mean(aic),
  1616. se_aic = sd(aic) / sqrt(n()),
  1617. .groups = "drop")
  1618. # Dynamic offset for the label (5% of y-range above baseline)
  1619. offset_acc <- diff(range(aic_long_acc_no_sm$aic)) * 0.05
  1620. ggplot() +
  1621. geom_hline(yintercept = baseline_acc,
  1622. linetype = "dashed", color = "coral", linewidth = 1) +
  1623. geom_jitter(data = aic_long_acc_no_sm,
  1624. aes(x = type, y = aic, color = type),
  1625. alpha = 0.3, width = 0.1, size = 2) +
  1626. geom_point(data = aic_se_acc,
  1627. aes(x = type, y = mean_aic, color = type),
  1628. size = 5) +
  1629. geom_errorbar(data = aic_se_acc,
  1630. aes(x = type, ymin = mean_aic - se_aic,
  1631. ymax = mean_aic + se_aic, color = type),
  1632. width = 0.15, linewidth = 1.2) +
  1633. annotate("text", x = 1.5, y = baseline_acc + offset_acc,
  1634. label = "Sensorimotor", color = "coral", size = 5) +
  1635. labs(x = NULL, y = "AIC (lower is better)",
  1636. title = "Predicting Accuracy") +
  1637. scale_color_manual(values = c("Distributional" = "gray60",
  1638. "Hybrid" = "steelblue")) +
  1639. theme_minimal() +
  1640. theme(text = element_text(size = 16), legend.position = "none")
  1641. # --- RT ---
  1642. aic_long_rt <- best_layers_rt %>%
  1643. select(Model, Layer, aic_sm, aic_bert, aic_combined) %>%
  1644. pivot_longer(cols = c(aic_sm, aic_bert, aic_combined),
  1645. names_to = "type", values_to = "aic") %>%
  1646. mutate(type = case_when(
  1647. type == "aic_sm" ~ "Sensorimotor",
  1648. type == "aic_bert" ~ "Distributional",
  1649. type == "aic_combined" ~ "Hybrid"
  1650. ),
  1651. type = factor(type, levels = c("Sensorimotor", "Distributional", "Hybrid")))
  1652. baseline_rt <- aic_long_rt %>% filter(type == "Sensorimotor") %>% pull(aic) %>% unique()
  1653. aic_long_rt_no_sm <- aic_long_rt %>% filter(type != "Sensorimotor")
  1654. aic_se_rt <- aic_long_rt_no_sm %>%
  1655. group_by(type) %>%
  1656. summarize(mean_aic = mean(aic),
  1657. se_aic = sd(aic) / sqrt(n()),
  1658. .groups = "drop")
  1659. offset_rt <- diff(range(aic_long_rt_no_sm$aic)) * 0.05
  1660. ggplot() +
  1661. geom_hline(yintercept = baseline_rt,
  1662. linetype = "dashed", color = "coral", linewidth = 1) +
  1663. geom_jitter(data = aic_long_rt_no_sm,
  1664. aes(x = type, y = aic, color = type),
  1665. alpha = 0.3, width = 0.1, size = 2) +
  1666. geom_point(data = aic_se_rt,
  1667. aes(x = type, y = mean_aic, color = type),
  1668. size = 5) +
  1669. geom_errorbar(data = aic_se_rt,
  1670. aes(x = type, ymin = mean_aic - se_aic,
  1671. ymax = mean_aic + se_aic, color = type),
  1672. width = 0.15, linewidth = 1.2) +
  1673. annotate("text", x = 1.5, y = baseline_rt + offset_rt,
  1674. label = "Sensorimotor", color = "coral", size = 5) +
  1675. labs(x = NULL, y = "AIC (lower is better)",
  1676. title = "Predicting RT") +
  1677. scale_color_manual(values = c("Distributional" = "gray60",
  1678. "Hybrid" = "steelblue")) +
  1679. theme_minimal() +
  1680. theme(text = element_text(size = 16), legend.position = "none")
  1681. ```
  1682. ## Part 3: Does SM account for sense boundaries?
  1683. ```{r}
  1684. # --- Compute AIC for each model/layer: same alone vs. same + continuous ---
  1685. tb_acc_q3 <- df_full_acc %>%
  1686. group_by(Model, Layer, n_params, Multilingual) %>%
  1687. summarize(
  1688. aic_same = AIC(lm(accuracy ~ same)),
  1689. aic_same_sm = AIC(lm(accuracy ~ same + sensorimotor_distance)),
  1690. aic_same_dist = AIC(lm(accuracy ~ same + Distance)),
  1691. aic_same_both = AIC(lm(accuracy ~ same + Distance + sensorimotor_distance)),
  1692. r2_same_dist = summary(lm(accuracy ~ same + Distance))$r.squared,
  1693. .groups = "drop"
  1694. ) %>%
  1695. mutate(
  1696. delta_aic_sm = aic_same - aic_same_sm, # positive = SM adds over same
  1697. delta_aic_dist = aic_same - aic_same_dist, # positive = Distributional adds over same
  1698. delta_aic_both = aic_same - aic_same_both # positive = both add over same
  1699. )
  1700. tb_rt_q3 <- df_full_rt %>%
  1701. group_by(Model, Layer, n_params, Multilingual) %>%
  1702. summarize(
  1703. aic_same = AIC(lm(mean_log_rt ~ same)),
  1704. aic_same_sm = AIC(lm(mean_log_rt ~ same + sensorimotor_distance)),
  1705. aic_same_dist = AIC(lm(mean_log_rt ~ same + Distance)),
  1706. aic_same_both = AIC(lm(mean_log_rt ~ same + Distance + sensorimotor_distance)),
  1707. r2_same_dist = summary(lm(mean_log_rt ~ same + Distance))$r.squared,
  1708. .groups = "drop"
  1709. ) %>%
  1710. mutate(
  1711. delta_aic_sm = aic_same - aic_same_sm,
  1712. delta_aic_dist = aic_same - aic_same_dist,
  1713. delta_aic_both = aic_same - aic_same_both
  1714. )
  1715. # --- Best layer per model (selected by best same + Distributional R²) ---
  1716. best_layers_acc_q3 <- tb_acc_q3 %>%
  1717. group_by(Model, n_params, Multilingual) %>%
  1718. slice_max(r2_same_dist, n = 1) %>%
  1719. ungroup()
  1720. best_layers_rt_q3 <- tb_rt_q3 %>%
  1721. group_by(Model, n_params, Multilingual) %>%
  1722. slice_max(r2_same_dist, n = 1) %>%
  1723. ungroup()
  1724. print(best_layers_acc_q3)
  1725. print(best_layers_rt_q3)
  1726. # --- Summary counts: how often does each augmentation add over same alone ---
  1727. best_layers_acc_q3 %>%
  1728. mutate(sm_adds = delta_aic_sm > 4,
  1729. dist_adds = delta_aic_dist > 4,
  1730. both_add = delta_aic_both > 4) %>%
  1731. summarise(across(c(sm_adds, dist_adds, both_add), mean))
  1732. best_layers_rt_q3 %>%
  1733. mutate(sm_adds = delta_aic_sm > 4,
  1734. dist_adds = delta_aic_dist > 4,
  1735. both_add = delta_aic_both > 4) %>%
  1736. summarise(across(c(sm_adds, dist_adds, both_add), mean))
  1737. make_aic_plot_q3 <- function(best_layers_df, title) {
  1738. aic_long <- best_layers_df %>%
  1739. select(Model, Layer, aic_same, aic_same_sm, aic_same_dist, aic_same_both) %>%
  1740. pivot_longer(cols = c(aic_same, aic_same_sm, aic_same_dist, aic_same_both),
  1741. names_to = "type", values_to = "aic") %>%
  1742. mutate(type = case_when(
  1743. type == "aic_same" ~ "Same only",
  1744. type == "aic_same_sm" ~ "Same + SM",
  1745. type == "aic_same_dist" ~ "Same + Distributional",
  1746. type == "aic_same_both" ~ "Same + Both"
  1747. ),
  1748. type = factor(type, levels = c("Same only", "Same + SM",
  1749. "Same + Distributional", "Same + Both")))
  1750. baseline <- aic_long %>% filter(type == "Same only") %>% pull(aic) %>% unique()
  1751. aic_long_no_base <- aic_long %>% filter(type != "Same only")
  1752. aic_se <- aic_long_no_base %>%
  1753. group_by(type) %>%
  1754. summarize(mean_aic = mean(aic),
  1755. se_aic = sd(aic) / sqrt(n()),
  1756. .groups = "drop")
  1757. offset <- diff(range(aic_long_no_base$aic)) * 0.05
  1758. ggplot() +
  1759. geom_hline(yintercept = baseline,
  1760. linetype = "dashed", color = "coral", linewidth = 1) +
  1761. geom_jitter(data = aic_long_no_base,
  1762. aes(x = type, y = aic, color = type),
  1763. alpha = 0.3, width = 0.1, size = 2) +
  1764. geom_point(data = aic_se,
  1765. aes(x = type, y = mean_aic, color = type),
  1766. size = 5) +
  1767. geom_errorbar(data = aic_se,
  1768. aes(x = type, ymin = mean_aic - se_aic,
  1769. ymax = mean_aic + se_aic, color = type),
  1770. width = 0.15, linewidth = 1.2) +
  1771. annotate("text", x = 2, y = baseline + offset,
  1772. label = "Same only", color = "coral", size = 5) +
  1773. labs(x = NULL, y = "AIC (lower is better)", title = title) +
  1774. scale_color_manual(values = c("Same + SM" = "coral",
  1775. "Same + Distributional" = "gray60",
  1776. "Same + Both" = "steelblue")) +
  1777. theme_minimal() +
  1778. theme(text = element_text(size = 16), legend.position = "none",
  1779. axis.text.x = element_text(angle = 20, hjust = 1))
  1780. }
  1781. make_aic_plot_q3(best_layers_acc_q3, "Predicting Accuracy")
  1782. make_aic_plot_q3(best_layers_rt_q3, "Predicting RT")
  1783. ```
  1784. # Comparison to Trott (2024) LLM Norms
  1785. Here, we conduct a more in-depth comparison to the GPT-4-generated norms from Trott (2024).
  1786. ## Load data
  1787. ```{r}
  1788. df_gpt_perception = read_csv("../../data/trott2024/cs_norms_perception_gpt-4.csv")
  1789. df_gpt_action = read_csv("../../data/trott2024/cs_norms_action_gpt-4.csv")
  1790. df_both_gpt = df_gpt_perception %>%
  1791. inner_join(df_gpt_action)
  1792. df_gpt_with_human = df_both_gpt %>%
  1793. inner_join(df_contextualized_meanings)
  1794. ### RAW-C With GPT and human SM distance
  1795. df_rawc_with_sm = read_csv("../../data/trott2024/rawc_with_sm.csv")
  1796. ```
  1797. ## Dimension-level correlation
  1798. ```{r gpt_corrs}
  1799. df_summ = df_gpt_with_human %>%
  1800. summarise(Vision = cor(Vision.M, Vision, method = "spearman"),
  1801. Hearing = cor(Hearing.M, Hearing, method = "spearman"),
  1802. Touch = cor(Touch.M, Touch, method = "spearman"),
  1803. Olfaction = cor(Olfaction.M, Olfaction, method = "spearman"),
  1804. Taste = cor(Taste.M, Taste, method = "spearman"),
  1805. Interoception = cor(Interoception.M, Interoception, method = "spearman"),
  1806. Hand_arm = cor(Hand_arm.M, Hand_arm, method = "spearman"),
  1807. Foot_leg = cor(Foot_leg.M, Foot_leg, method = "spearman"),
  1808. Head = cor(Head.M, Head, method = "spearman"),
  1809. Torso = cor(Torso.M, Torso, method = "spearman"),
  1810. Mouth_throat = cor(Mouth_throat.M, Mouth_throat, method = "spearman"))
  1811. df_summ
  1812. df_long = df_summ %>%
  1813. pivot_longer(everything(), names_to = "Factor", values_to = "Correlation")
  1814. df_long %>%
  1815. ggplot(aes(x = reorder(Factor, Correlation), y = Correlation)) +
  1816. geom_bar(stat = "identity") +
  1817. labs(x = "Factor", y = "Correlation") +
  1818. theme_minimal()
  1819. ```
  1820. ## Compare correlation matrices
  1821. ```{r gpt_corrs_comparison}
  1822. library(tidyverse)
  1823. library(ggcorrplot)
  1824. library(patchwork)
  1825. dim_order <- c("Vision", "Hearing", "Olfaction", "Taste", "Interoception",
  1826. "Touch", "Mouth_throat", "Head", "Torso", "Hand_arm", "Foot_leg")
  1827. # Human matrix: .M columns
  1828. human_cols <- df_gpt_with_human %>%
  1829. select(all_of(paste0(dim_order, ".M"))) %>%
  1830. rename_with(~ str_remove(., "\\.M"))
  1831. cors_human <- cor(human_cols, use = "pairwise.complete.obs")
  1832. # GPT matrix: bare-named columns
  1833. gpt_cols <- df_gpt_with_human %>%
  1834. select(all_of(dim_order))
  1835. cors_gpt <- cor(gpt_cols, use = "pairwise.complete.obs")
  1836. # Difference (GPT - Human)
  1837. cors_diff <- cors_gpt - cors_human
  1838. ggcorrplot(cors_human, hc.order = FALSE, type = "upper",
  1839. lab = TRUE, lab_size = 2.5) +
  1840. ggtitle("Human") +
  1841. theme(axis.text.x = element_text(size = 10, angle = 45, hjust = 1))
  1842. ggcorrplot(cors_gpt, hc.order = FALSE, type = "upper",
  1843. lab = TRUE, lab_size = 2.5) +
  1844. ggtitle("GPT-4") +
  1845. theme(axis.text.x = element_text(size = 10, angle = 45, hjust = 1))
  1846. ggcorrplot(cors_diff, hc.order = FALSE, type = "upper",
  1847. lab = TRUE, lab_size = 2.5,
  1848. colors = c("#2166AC", "white", "#B2182B")) +
  1849. ggtitle("GPT − Human") +
  1850. theme(axis.text.x = element_text(size = 10, angle = 45, hjust = 1))
  1851. ### Statistical comparison
  1852. library(vegan)
  1853. # Mantel test
  1854. mantel_result <- mantel(cors_human, cors_gpt,
  1855. method = "pearson", # or "spearman"
  1856. permutations = 9999)
  1857. mantel_result
  1858. mantel_result <- mantel(cors_human, cors_gpt,
  1859. method = "spearman", # or "spearman"
  1860. permutations = 9999)
  1861. mantel_result
  1862. ## Or just straight-up correlation
  1863. upper_human <- cors_human[upper.tri(cors_human)]
  1864. upper_gpt <- cors_gpt[upper.tri(cors_gpt)]
  1865. # Correlation between the two
  1866. cor(upper_human, upper_gpt, method = "pearson")
  1867. cor(upper_human, upper_gpt, method = "spearman")
  1868. ```
  1869. ## Substitution analyses
  1870. ### Homonymy vs. Polysemy
  1871. Here, we ask whether homonyms show bigger sensorimotor distance for GPT-generated norms as well.
  1872. ```{r}
  1873. model_full = lmer(data = df_rawc_with_sm,
  1874. sm_gpt ~ same +
  1875. (1 + same | word),
  1876. control=lmerControl(optimizer="bobyqa"),
  1877. REML = FALSE)
  1878. model_reduced = lmer(data = df_rawc_with_sm,
  1879. sm_gpt ~ # same +
  1880. (1 + same | word),
  1881. control=lmerControl(optimizer="bobyqa"),
  1882. REML = FALSE)
  1883. summary(model_full)
  1884. anova(model_full, model_reduced)
  1885. model_full_at = lmer(data = df_rawc_with_sm,
  1886. sm_gpt ~ same * ambiguity_type +
  1887. (1 + same | word),
  1888. control=lmerControl(optimizer="bobyqa"),
  1889. REML = FALSE)
  1890. model_reduced_at = lmer(data = df_rawc_with_sm,
  1891. sm_gpt ~ same + ambiguity_type +
  1892. (1 + same | word),
  1893. control=lmerControl(optimizer="bobyqa"),
  1894. REML = FALSE)
  1895. summary(model_full_at)
  1896. anova(model_full_at, model_reduced_at)
  1897. df_rawc_with_sm %>%
  1898. group_by(same, ambiguity_type) %>%
  1899. summarise(mean_gpt = mean(sm_gpt),
  1900. sd_gpt = sd(sm_gpt),
  1901. mean_human = mean(sm_human),
  1902. sd_human = sd(sm_human))
  1903. df_rawc_with_sm = df_rawc_with_sm %>%
  1904. mutate(Same = case_when(
  1905. same == TRUE ~ "Same Sense",
  1906. same == FALSE ~ "Different Sense"
  1907. ))
  1908. df_rawc_with_sm %>%
  1909. ggplot(aes(x = sm_gpt,
  1910. y = ambiguity_type,
  1911. fill = Same)) +
  1912. geom_density_ridges2(aes(height = ..density..),
  1913. color = NA,
  1914. alpha = 0.5,
  1915. scale=0.85,
  1916. stat="density") +
  1917. labs(x = "Sensorimotor Distance (GPT-4)",
  1918. y = NULL,
  1919. fill = NULL) +
  1920. scale_fill_manual(
  1921. values = c("Same Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[1],
  1922. "Different Sense" = viridisLite::viridis(2, option = "mako", begin = 0.8, end = 0.15)[2])
  1923. ) +
  1924. theme_minimal() +
  1925. theme(text = element_text(size = 20))
  1926. ```
  1927. ### Dominance
  1928. We also replicate the dominance analysis.
  1929. ```{r gpt_sm_strength}
  1930. df_dom_gpt = df_gpt_with_human %>%
  1931. inner_join(df_dominance_individual)
  1932. df_dom_gpt <- df_dom_gpt %>%
  1933. inner_join(df_lancaster)
  1934. nrow(df_dom_gpt)
  1935. df_dom_gpt = df_dom_gpt %>%
  1936. rowwise() %>%
  1937. mutate(max_strength_gpt = max(
  1938. c(
  1939. ## Modalities
  1940. Vision,
  1941. Hearing,
  1942. Olfaction,
  1943. Touch,
  1944. Taste,
  1945. Interoception,
  1946. ## Effectors
  1947. Head,
  1948. Mouth_throat,
  1949. Torso,
  1950. Hand_arm,
  1951. Foot_leg
  1952. )
  1953. ),
  1954. minkowski3_strength_gpt = sum(c(Vision,
  1955. Hearing,
  1956. Olfaction,
  1957. Touch,
  1958. Taste,
  1959. Interoception,
  1960. ## Effectors
  1961. Head,
  1962. Mouth_throat,
  1963. Torso,
  1964. Hand_arm,
  1965. Foot_leg)^3)^(1/3),
  1966. max_strength = max(
  1967. c(
  1968. ## Modalities
  1969. Vision.M,
  1970. Hearing.M,
  1971. Olfaction.M,
  1972. Touch.M,
  1973. Taste.M,
  1974. Interoception.M,
  1975. ## Effectors
  1976. Head.M,
  1977. Mouth_throat.M,
  1978. Torso.M,
  1979. Hand_arm.M,
  1980. Foot_leg.M
  1981. )
  1982. ),
  1983. minkowski3_strength = sum(c(Vision.M, Hearing.M, Olfaction.M, Touch.M, Taste.M,
  1984. Interoception.M, Head.M, Mouth_throat.M, Torso.M,
  1985. Hand_arm.M, Foot_leg.M)^3)^(1/3)
  1986. ) %>%
  1987. ungroup()
  1988. cor.test(df_dom_gpt$max_strength_gpt, df_dom_gpt$max_strength)
  1989. cor.test(df_dom_gpt$minkowski3_strength_gpt, df_dom_gpt$minkowski3_strength)
  1990. df_dom_gpt %>%
  1991. ggplot(aes(x = minkowski3_strength_gpt,
  1992. y = minkowski3_strength)) +
  1993. geom_point(alpha = .5) +
  1994. geom_smooth(method = "lm") +
  1995. labs(x = "Minkowski-3 Strength (GPT-4)",
  1996. y = "Minkowski-3 Strength (Human)") +
  1997. theme_minimal() +
  1998. theme(text = element_text(size = 16), legend.position = "none")
  1999. ```
  2000. ### Does GPT sensorimotor strength predict dominance?
  2001. ```{r gpt_dominance}
  2002. df_dom_gpt %>%
  2003. ggplot(aes(x = minkowski3_strength,
  2004. y = dominance)) +
  2005. geom_point(alpha = .4) +
  2006. geom_smooth(method = "lm") +
  2007. labs(x = "Minkowski-3 Sensorimotor Strength (GPT-4)",
  2008. y = "Dominance") +
  2009. theme_minimal() +
  2010. theme(text = element_text(size=16))
  2011. model_full = lmer(data = df_dom_gpt,
  2012. dominance ~
  2013. max_strength_gpt +
  2014. Max_strength.sensorimotor + Minkowski3.sensorimotor +
  2015. (1 | word),
  2016. REML = FALSE)
  2017. model_reduced = lmer(data = df_dom_gpt,
  2018. dominance ~
  2019. # max_strength_gpt +
  2020. Max_strength.sensorimotor + Minkowski3.sensorimotor +
  2021. (1 | word),
  2022. REML = FALSE)
  2023. summary(model_full)
  2024. anova(model_full, model_reduced)
  2025. model_full = lmer(data = df_dom_gpt,
  2026. dominance ~
  2027. minkowski3_strength_gpt +
  2028. Max_strength.sensorimotor + Minkowski3.sensorimotor +
  2029. (1 | word),
  2030. REML = FALSE)
  2031. model_reduced = lmer(data = df_dom_gpt,
  2032. dominance ~
  2033. # minkowski3_strength_gpt +
  2034. Max_strength.sensorimotor + Minkowski3.sensorimotor +
  2035. (1 | word),
  2036. REML = FALSE)
  2037. summary(model_full)
  2038. anova(model_full, model_reduced)
  2039. ```

contextualized_norms_analysis.Rmd at commit 7f87ad2, no license · at the source

Overview

Authors: Sean Trott1,2, Benjamin Bergen2
ORCID iDs: Sean Trott
  1. Rutgers University-Newark, Newark, NJ USA
  2. UC San Diego, La Jolla, CA USA
Journal: Behavior research methods, volume 58, issue 9, article 264
Dates: received 7 February 2026; accepted 6 July 2026; published online 12 August 2026; in print 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.3758/s13428-026-03134-6 · PMID 42587187 · PMCID PMC13469423 · OpenAlex W4221144368
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: human (organism)
Methods: Statistics, Machine learning, Connectivity
MeSH: Language*, Linguistics*, Psycholinguistics*, Female, Humans, Judgment, Semantics (* major topic)
Topic: Action Observation and Synchronization (Social Psychology, Psychology), according to OpenAlex
Citations: not cited yet (Europe PMC); 111 references in the paper

Abstract

Embodied theories of language emphasize the role of sensorimotor experience in linguistic knowledge. Central to testing these theories is the creation of large datasets of linguistic norms, which contain judgments about a word’s sensorimotor associations and can be used to predict human behavioral or brain data – sometimes in contrast to competing variables, such as those derived from distributional language models. Yet many of these datasets contain judgments about words in isolation, despite the fact that most words are ambiguous, making it difficult to determine which meaning of a word is characterized by its rating (e.g., “wooden table” vs. “data table”). In the current work, we introduce a new lexical resource (directly inspired by the Lancaster sensorimotor for 112 English words, each rated in four different contexts (448 sentences total). We demonstrate: first, that these ratings encode overlapping but distinct information from the Lancaster sensorimotor norms; second, that decontextualized ratings likely reflect the more dominant meaning of ambiguous words; third, that homonyms have more distinct sensorimotor profiles than polysemes; fourth, that the contextualized sensorimotor distance between two uses of an ambiguous word predicts human judgments about semantic relatedness; and fifth, that ratings derived from GPT-4 align reasonably well with human judgments. We conclude by suggesting that contextualized ratings like these can be used both to inform competing theories of semantic representations and also to evaluate or “probe” the ability of LMs to recover sensorimotor information.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 6 matches between paragraphs and lines of code.

seantrott/cs_norms

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 7f87ad29ff1280087f995f7e691a4bb0ab12ffcb, 1 June 2026
Languages: Python (2), R (1)
Size: 42 files, 3 scripts
Software Heritage: archived
Found in: “Code Availability”
Holds: README, 1 notebook
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (2 files), pandas (2 files), SciPy (2 files), broom (1 file), lme4 (1 file), lmerTest (1 file), patchwork (1 file), tidyverse (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
4 files

Code Availability

Analysis code can be found on GitHub: https://github.com/seantrott/cs_norms.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 3 scripts, each with its path and the digest of its content;
  • 6 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data Availability Statement

The contextualized sensorimotor norms are publicly available on GitHub (https://github.com/seantrott/cs_norms) and HuggingFace (https://huggingface.co/datasets/seantrott/cs_norms).

Analysis code can be found on GitHub: https://github.com/seantrott/cs_norms.

Reproduced under the paper's license (CC BY), from the paper cited above.

Open Practices Statement

All code and data can be found in a publicly available repository (see above).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 3, 28 September 2026

  • Publisher: n/a → Springer Science+Business Media

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 2 authors, 7 MeSH terms, 59 references.

Cite

This paper

Trott, S., & Bergen, B. (2026). Contextualized sensorimotor norms: Multi-dimensional measures of sensorimotor strength for ambiguous English words, in context. Behavior research methods, 58(9), 264. https://doi.org/10.3758/s13428-026-03134-6

BibTeX

@article{trott2026contextualized,
author = {Trott, Sean and Bergen, Benjamin},
title = {{Contextualized sensorimotor norms: Multi-dimensional measures of sensorimotor strength for ambiguous English words, in context}},
journal = {Behavior research methods},
year = {2026},
month = aug,
volume = {58},
number = {9},
pages = {264},
publisher = {Springer Science+Business Media},
issn = {1554-351X},
doi = {10.3758/s13428-026-03134-6},
url = {https://doi.org/10.3758/s13428-026-03134-6},
pmid = {42587187},
pmcid = {PMC13469423}
}

RIS

TY - JOUR
AU - Trott, Sean
AU - Bergen, Benjamin
TI - Contextualized sensorimotor norms: Multi-dimensional measures of sensorimotor strength for ambiguous English words, in context
T2 - Behavior research methods
J2 - Behav Res Methods
PY - 2026
DA - 2026/08/12
VL - 58
IS - 9
SP - 264
SN - 1554-351X
PB - Springer Science+Business Media
DO - 10.3758/s13428-026-03134-6
UR - https://doi.org/10.3758/s13428-026-03134-6
LA - en
ER -

CSL-JSON

{
"id": "10.3758/s13428-026-03134-6",
"type": "article-journal",
"title": "Contextualized sensorimotor norms: Multi-dimensional measures of sensorimotor strength for ambiguous English words, in context",
"container-title": "Behavior research methods",
"author": [
{
"family": "Trott",
"given": "Sean"
},
{
"family": "Bergen",
"given": "Benjamin"
}
],
"container-title-short": "Behav Res Methods",
"volume": "58",
"issue": "9",
"page": "264",
"DOI": "10.3758/s13428-026-03134-6",
"PMID": "42587187",
"PMCID": "PMC13469423",
"ISSN": "1554-351X",
"publisher": "Springer Science+Business Media",
"URL": "https://doi.org/10.3758/s13428-026-03134-6",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
12
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1073/pnas.2603114123 [code]
The human hippocampus can pattern separate memories by meaning.
Journal: Proceedings of the National Academy of Sciences of the United States of America
In common: broom, lmerTest, lme4, 5 other tools
[2] doi:10.1016/j.dcn.2026.101765 [code]
Fusiform face area development correlates with development in higher-order social brain regions.
Journal: Developmental cognitive neuroscience
In common: broom, lmerTest, lme4, 5 other tools
[3] doi:10.1162/imag.a.1235 [code]
Intracranial volume: To adjust or not to adjust? It is not a matter of if, but how.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: broom, lmerTest, lme4, 5 other tools
[4] doi:10.64898/2026.03.09.710596 [code]
Infant gut microbiomes contribute to metabolic states that impact brain function
Journal: bioRxiv (preprint)
In common: broom, lmerTest, lme4, 4 other tools, 1 reference
[5] doi:10.1162/imag.a.1321 [code]
Phase similarity between similar objects indicates representational merging across retrieval training but not sleep.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: broom, lmerTest, lme4, 4 other tools, 1 reference
[6] doi:10.1093/cercor/bhag132 [code]
Spatiotemporal white-matter development across early childhood.
Journal: Cerebral cortex (New York, N.Y. : 1991)
In common: broom, lmerTest, lme4, 4 other tools
[7] doi:10.1038/s41467-026-71264-8 [code]
Habitual coffee intake shapes the gut microbiome and modifies host physiology and cognition.
Journal: Nature communications
In common: broom, lmerTest, lme4, 4 other tools
[8] doi:10.1016/j.celrep.2026.117505 [code]
Impaired spatial coding and neuronal hyperactivity in the medial entorhinal cortex of aged APP knock-in mice.
Journal: Cell reports
In common: broom, lmerTest, lme4, 4 other tools
[9] doi:10.3390/ijms27135713 [code]
Chronic Administration of Marinobufagenin in Mice Causes Hyperlocomotion and Decrease in Anxiety by Altering Monoamine Turnover Unaccompanied by Motor Deficits or Oxidative Stress.
Journal: International journal of molecular sciences
In common: broom, lmerTest, lme4, 4 other tools
[10] doi:10.1093/ageing/afag263 [code]
Cardiometabolic medication exposures and cognitive outcomes in Alzheimer's disease.
Journal: Age and ageing
In common: broom, lmerTest, lme4, 4 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.