OSCR

Prior cocaine use disrupts identification of hidden states by single units and neural ensembles in orbitofrontal cortex.

Code ↔ Paper

7 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 7 matches
  1. [1] § Methods › Quantification and statistical analyses › Classification analyses ↔ !!! osf code/supporting_scripts/ndt.1.0.4/ndt.1.0.4_exported/classifiers/@libsvm_CL/libsvm_CL.m, lines 1–129 · score 0.61 · Support Vector Machine, classification, linear, dimensional, accuracy, matrix
  2. [2] § Methods › Quantification and statistical analyses › Classification analyses ↔ !!! osf code/supporting_scripts/ndt.1.0.4/ndt.1.0.4_exported/cross_validators/@standard_resample_CV/standard_resample_CV.m, lines 1–60 · score 0.59 · cross validation procedure, decoding accuracy, classification, temporal, matrix, train
  3. [3] § Methods › Quantification and statistical analyses › TCA analysis ↔ !!! osf code/TCA/python/tca_fig8.py, lines 62–114 · score 0.56 · reconstruction error, optimization, TCA, fit, tensor, components
  4. [4] § Methods › Quantification and statistical analyses › TCA analysis ↔ !!! osf code/data_file_formatting/get_psth_figure8_cocaine.m, lines 302–362 · score 0.56 · postTrial1, postTrial2, preTrial, firing, Unpoke, Odor
  5. [5] § Methods › Quantification and statistical analyses › Task events and peri-event spike train analysis ↔ !!! osf code/supporting_scripts/ndt.1.0.4/ndt.1.0.4_exported/datasources/@basic_DS/basic_DS.m, lines 174–220 · score 0.55 · random selection, firing rates, spike, bin, neuron
  6. [6] § Methods › Quantification and statistical analyses › Cross-sequence decoding ↔ !!! osf code/supporting_scripts/ndt.1.0.4/ndt.1.0.4_exported/datasources/@basic_DS/basic_DS.m, lines 127–171 · score 0.52 · cross validation, Decoding accuracy, iteration, population, cells, trained
  7. [7] § Results ↔ !!! osf code/behav_analysis/plot_behav_figure8_cocaine3.m, lines 436–483 · score 0.51 · absolute poke latency, Error bars, locations, SEMs, sucrose, cocaine

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

MATLAB · 735 lines · 41 KB · no license · 2 matches

  1. classdef basic_DS < handle
  2. % basic_DS implements the basic functions of a
  3. % datasource (DS) object, namely, it takes binned data and labels,
  4. % and through the get_data method, the object returns k
  5. % leave-one-fold-out cross-validation splits of the data which can subsequently
  6. % be used to train and test a classifier. The data in the population vectors is
  7. % randomly selected from the larger binned data that is passed to the constructor
  8. % of this object. This object can create both
  9. % pseduo-populations (i.e., populations vector in which the recordings were
  10. % made on independent sessions but are treated as if they were recorded simultaneously)
  11. % and simultaneously populations in which neurons that were recorded together
  12. % always appear together in population vectors.
  13. %
  14. % Like all DS objects, basic_DS implements the method get_data, which has the following form:
  15. %
  16. % [XTr_all_time_cv YTr_all XTe_all_time_cv YTe_all] = get_data(ds); where:
  17. %
  18. % a. XTr_all_time_cv{iTime}{iCV} = [num_features x num_training_points] is a
  19. % cell array that has the training data for all times and cross-validation splits
  20. % b. YTr_all = [num_training_point x 1] a vector of the training labels (the same training labels are used at all times and CV splits)
  21. % c. XTe_all_time_cv{iTime}{iCV} = [num_features x num_test_points] is a
  22. % cell array that has the test data for all times and cross-validation splits;
  23. % d. YTe_all = [num_test_point x 1] a vector has the test labels (the same test labels are used at all times and CV splits)
  24. %
  25. %
  26. % The constructor for this object has the form:
  27. %
  28. % ds = basic_DS(binned_data_name, specific_binned_label_name, num_cv_splits, load_data_as_spike_counts), where:
  29. %
  30. % a. binned_data_name: is string that has the name of a file that has data in binned-format, or is a cell array of binned-format binned_data
  31. % b. specific_binned_labels_name: is a string containing a specific binned-format label name, or is a cell array/vector containing
  32. % the specific binned names (i.e., binned_labels.specific_binned_label_name)
  33. % c. num_cv_splits = is a scalar indicating how many cross-validation splits there should be
  34. % d. load_data_as_spike_counts: an optional flag that can be set that will cause the data to be converted to spike counts if set to an integer rather than 0
  35. % (the create_binned_data_from_raster_data function saves data as firing rates by default). This flag is useful
  36. % when using the Poison Naive Bayes classifier that needs spike counts rather than firing rates. If this flag is not set, the default behavior
  37. % is to use firing rates.
  38. %
  39. %
  40. % The basic_DS also has the following properties that can be set:
  41. %
  42. % 1. create_simultaneously_recorded_populations (default = 0). If the data from all sites
  43. % was recorded simultaneously, then setting this variable to 1 causes the
  44. % function to return simultaneous populations rather than pseudo-populations
  45. % (for this to work all sites in 'the_data' must have the trials in the same order).
  46. % If this variable is set to 2, then the training set is pseudo-populations and the
  47. % test set is simultaneous populations. This allows one to estimate I_diag, as
  48. % described by Averbeck, Latham and Pouget in 'Neural correlations, population coding
  49. % and computation', Nature Neuroscience, May, 2006. I_diag is a measure that gives a
  50. % sense of whether training on pseudo-populations leads to a the same decision rule as
  51. % when training on simultaneous populations.
  52. %
  53. % 2. sample_sites_with_replacement (default = 0). This variable specifies whether
  54. % the sites should be sample with replacement - i.e., if the data is
  55. % sampled with replacement, then some sites will be repeated within a single
  56. % population vector. This allows one to do a bootstrap estimate of variance
  57. % of the results if different sites from a larger population had been selected
  58. % while also ensuring that there is no overlapping data between the training
  59. % and test sets.
  60. %
  61. % 3. num_times_to_repeat_each_label_per_cv_split (default = 1). This variable
  62. % specifies how many times each label should appear in each cross-validation split.
  63. % For example, if this value is set to k, this means that there will be k
  64. % population vectors from each class in each test set, and there will be
  65. % k * (num_cv_splits - 1) population vectors for each class in each training set split.
  66. %
  67. % 4. label_names_to_use (default = [] meaning all unique label names in the_labels are used).
  68. % This specifies which labels names (or numbers) to use, out of the unique label
  69. % names that are present in the the_labels cell array. If only a subset of labels are listed,
  70. % then only population vectors that have the specified labels will be returned.
  71. %
  72. % 5. num_resample_sites (default = -1, which means use all sites). This variable specifies
  73. % how many sites should be randomly selected each time the get_data method is called.
  74. % For example, suppose length(the_data) = n, and num_resample_sites = k, then each
  75. % time get_data is called, k of the n sites would randomly be selected to be included
  76. % as features in the population vector.
  77. %
  78. % 6. sites_to_use (default = -1, which means select features from all sites). This
  79. % variable allows one to only choose features from the sites listed in this vector
  80. % (i.e., features will only be randomly selected from the sites listed in this vector).
  81. %
  82. % 7. sites_to_exclude (default = [], which means do not exclude any sites). This allows
  83. % one to not select features from particular sites (i.e., features will NOT be
  84. % selected from the sites listed in this vector).
  85. %
  86. % 8. time_periods_to_get_data_from (default = [], which means create one feature
  87. % for all times that are present in the_data{iSite} matrix). This variable
  88. % can be set to a cell array that contains vectors that specify which time bins
  89. % to use as features from the_data. For examples, if time_periods_to_get_data_from = {[2 3], [4 5], [10 11]}
  90. % then there will be three time periods for XTr_all_time_cv and XTe_all_time_cv
  91. % (i.e., length(XTr_all_time_cv) = 3), and the population vectors for the
  92. % time period will have 2 * num_resample_sites features, with the population
  93. % vector for the first time period having data from each resample site from times
  94. % 2 and 3 in the_data{iSite} matrix, etc..
  95. %
  96. % 9. randomly_shuffle_labels_before_running (default = 0). If this variable is set to one
  97. % then the labels are randomly shuffled prior to the get_data method being called (thus all calls
  98. % to get_data return the same randomly shuffled labels). This method is useful for creating a
  99. % null distribution to test whether a decoding result is above what one would expect by chance.
  100. %
  101. %
  102. % This object also has two addition method which are:
  103. %
  104. % 1. the_properties = get_DS_properties(ds)
  105. % This method returns the main property values of the datasource.
  106. %
  107. % 2. ds = set_specific_sites_to_use(ds, curr_resample_sites_to_use)
  108. % This method causes the get_data to use specific sites rather than
  109. % choosing sites randomly. This method should really only be used by
  110. % other datasources that are extending the functionality of basic_DS.
  111. %
  112. %
  113. % Note: this class is a subclass of the handle class, meaning that when this object is created a
  114. % reference to the object is returned. Thus when fields of the object are changed a copy of the
  115. % object does not need to be returned (by default matlab objects are passed by value). The
  116. % advantage of having this object inherit from the handle class is that if the object changes its
  117. % state within a method, a copy of the object does not need to be returned (this is particularly
  118. % useful for the randomly_permute_labels_before_running method so that the labels can be randomly
  119. % shuffled once prior to the get_data method being called, and the same shuffled labels will
  120. % be used throughout all subsequent calls to get_data, allowing one to create a full null distribution
  121. % by running the code multiple times).
  122. %
  123. %==========================================================================
  124. % This code is part of the Neural Decoding Toolbox.
  125. % Copyright (C) 2011 by Ethan Meyers ([email hidden])
  126. %
  127. % This program is free software: you can redistribute it and/or modify
  128. % it under the terms of the GNU General Public License as published by
  129. % the Free Software Foundation, either version 3 of the License, or
  130. % (at your option) any later version.
  131. %
  132. % This program is distributed in the hope that it will be useful,
  133. % but WITHOUT ANY WARRANTY; without even the implied warranty of
  134. % MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
  135. % GNU General Public License for more details.
  136. %
  137. % You should have received a copy of the GNU General Public License
  138. % along with this program. If not, see <http://www.gnu.org/licenses/>.
  139. %==========================================================================
  140. properties
  141. the_labels % a cell array that contains vectors of labels that specify what occurred during each trial for all neurons the_data cell array
  142. num_cv_splits % how many cross-validation splits there should be
  143. num_times_to_repeat_each_label_per_cv_split = 1; % how many of each unique label should be in each CV block
  144. label_names_to_use = []; % which set of labels names should be used (or which numbers should be used if the_labels{iSite} is a vector of numbers)
  145. num_resample_sites = -1; % how many sites should be used for each resample iteration - must be less than length(the_data)
  146. sample_sites_with_replacement = 0; % specify whether to sample neurons with replacement - if they are sampled with replacement, then some features will be repeated within a single data vector
  147. % which reduces the number of neurons used (and thus usually lowers the decoding accuracy). It should be noted that all data duplication appears in the same CV trials
  148. % so there is no contamination with having repeatd data in different CV trials
  149. create_simultaneously_recorded_populations = 0; % to use pseudo-populations or simultaneous populations (2 => that the training set is pseudo and test is simultaneous)
  150. sites_to_use = -1; % a list of indices of which features (e.g., sites/neurons) to use in the the_data cell array
  151. sites_to_exclude = []; % a list of features that should explicitly be excluded
  152. time_periods_to_get_data_from = []; % a cell array containing vectors that specify which time bins to use from the_data
  153. % randomly shuffles the labels prior to the get_data method being called - which is useful for creating one point in a null distribution to check if decoding results are above what is expected by change.
  154. randomly_shuffle_labels_before_running = 0;
  155. binned_site_info % a variable that can contain the binned_site_info (this will be automatically set the data in the constructor is loaded from a string that has a file name of data in binned-format.)
  156. % this information is not used by the datasource, but it is returned by the get_properties method.
  157. % excluding this option for now
  158. % set these if you want the datasource to return different random selection of labels (and data from those trials) each time
  159. % use_random_subset_of_k_labels_each_time_data_is_retrieved = -1; % if this is set to k > 1, then a random subset of labels k labels will be chosen each resample iteration
  160. % 10. use_random_subset_of_k_labels_each_time_data_is_retrieved (default = -1). If this
  161. % variable is set to k > 1, then a random subset of labels k labels will be chosen each
  162. % time get_data is run (i.e., the different resample runs will use a different subset
  163. % of k labels each time the get_data method is called).
  164. end
  165. properties (GetAccess = 'public', SetAccess = 'private')
  166. the_data % a cell array that contains all the data binned data in the format the_data{iNeuron}[num_trials x num_time_bins]
  167. curr_resample_sites_to_use = -1; % This specified which features should be used on a given random resample iteration.
  168. % This should really be randomly selected each time the get_data is called, but some rare cases
  169. % it is useful to set to some specific features. The method set_specific_sites_to_use that allows this variable
  170. % to be set by outside calls.
  171. initialized = 0;
  172. label_names_to_label_numbers_mapping = [];
  173. data_loaded_as_spike_counts; % records whether the data was loaded as spike counts (rather than firing rates)
  174. end
  175. methods
  176. %function ds = basic_DS
  177. %end
  178. % the constructor
  179. function ds = basic_DS(binned_data_name, specific_binned_label_name, num_cv_splits, load_data_as_spike_counts)
  180. if nargin < 4
  181. load_data_as_spike_counts = 0;
  182. end
  183. if load_data_as_spike_counts > 0
  184. if ~isstr(binned_data_name)
  185. error('If the argument load_data_as_spike_counts is set to a value greater than 0, binned_data_name must be a string listing the name of a file that has data in binned format')
  186. end
  187. [binned_data_spike_counts binned_labels binned_site_info] = load_binned_data_and_convert_firing_rates_to_spike_counts(binned_data_name);
  188. ds.the_data = binned_data_spike_counts;
  189. ds.binned_site_info = binned_site_info;
  190. elseif isstr(binned_data_name)
  191. load(binned_data_name)
  192. ds.the_data = binned_data;
  193. ds.binned_site_info = binned_site_info;
  194. else
  195. ds.the_data = binned_data_name;
  196. end
  197. if isstr(specific_binned_label_name)
  198. ds.the_labels = eval(['binned_labels.' specific_binned_label_name]);
  199. else
  200. ds.the_labels = specific_binned_label_name;
  201. end
  202. ds.num_cv_splits = num_cv_splits;
  203. ds.data_loaded_as_spike_counts = load_data_as_spike_counts; % might as well save this information too
  204. end
  205. % This method allows one to set exact prespecified sites to get data from
  206. % rather than randomly selecting a set of sites from the larger population (it should rarely be used)
  207. function ds = set_specific_sites_to_use(ds, curr_resample_sites_to_use)
  208. ds.curr_resample_sites_to_use = curr_resample_sites_to_use;
  209. end
  210. % This method returns the main property values of the datasource (could be useful for saving what parameters were used)
  211. function the_properties = get_DS_properties(ds)
  212. the_properties.num_cv_splits = ds.num_cv_splits;
  213. the_properties.num_times_to_repeat_each_label_per_cv_split = ds.num_times_to_repeat_each_label_per_cv_split;
  214. the_properties.sample_sites_with_replacement = ds.sample_sites_with_replacement;
  215. the_properties.num_resample_sites = ds.num_resample_sites;
  216. the_properties.create_simultaneously_recorded_populations = ds.create_simultaneously_recorded_populations;
  217. the_properties.sites_to_use = ds.sites_to_use;
  218. the_properties.sites_to_exclude = ds.sites_to_exclude;
  219. the_properties.time_periods_to_get_data_from = ds.time_periods_to_get_data_from;
  220. the_properties.randomly_shuffle_labels_before_running = ds.randomly_shuffle_labels_before_running;
  221. the_properties.binned_site_info = ds.binned_site_info;
  222. the_properties.data_loaded_as_spike_counts = ds.data_loaded_as_spike_counts;
  223. %the_properties.use_random_subset_of_k_labels_each_time_data_is_retrieved = ds.use_random_subset_of_k_labels_each_time_data_is_retrieved;
  224. % if haven't converted to ds.label_names_to_use to strings yet (b/c get_data has not yet been called), or if the_labels is numbers, just return input set by user
  225. if iscell(ds.label_names_to_use) || isempty(ds.label_names_to_label_numbers_mapping)
  226. the_properties.label_names_to_use = ds.label_names_to_use;
  227. else % if code has already converted ds.label_names_to_use from strings to numbers convert them back to strings
  228. for iName = 1:length(ds.label_names_to_use)
  229. the_properties.label_names_to_use{iName} = ds.label_names_to_label_numbers_mapping{ds.label_names_to_use(iName)}; % is cell array so won't work
  230. end
  231. end
  232. end
  233. function [XTr_all_time_cv YTr_all XTe_all_time_cv YTe_all] = get_data(ds)
  234. % The main DS function that returns training and test population vectors. The outputs of this function are:
  235. %
  236. % 1. XTr_all_time_cv{iTime}{iCV} = [num_features x num_training_points] is a
  237. % cell array that has the training data for all times and cross-validation splits
  238. %
  239. % 2. YTr_all{iTime} = [num_training_point x 1] has the training labels
  240. %
  241. % 3. XTe_all_time_cv{iTime}{iCV} = [num_features x num_test_points] is a
  242. % cell array that has the test data for all times and cross-validation splits
  243. %
  244. % 4. YTe_all{iTime} = [num_test_point x 1] has the test labels
  245. %
  246. % initialize variables the first time ds.get_data is called
  247. if ds.initialized == 0
  248. disp('initializing basic_DS.get_data')
  249. % if the_labels is a cell array of strings, convert the_labels into a vector of numbers
  250. if iscell(ds.the_labels{1}) % just checking the first site (assuming it will be the same for all other sites)
  251. ignore_case_of_strings = 0; % for now, always respect the case of the strings used in the labels
  252. % % [specific_binned_labels_as_numbers string_to_number_mapping] = convert_label_strings_into_numbers(ds.the_labels, ignore_case_of_strings, ds.label_string_names_to_use);
  253. % doing it this way causes a consistent mapping from label strings to label numbers, and then the strings to be used are selected through setting ds.label_names_to_use
  254. [specific_binned_labels_as_numbers string_to_number_mapping] = convert_label_strings_into_numbers(ds.the_labels, ignore_case_of_strings);
  255. ds.the_labels = specific_binned_labels_as_numbers;
  256. ds.label_names_to_label_numbers_mapping = string_to_number_mapping;
  257. if ~isempty(ds.label_names_to_use)
  258. label_numbers_used = find(ismember(string_to_number_mapping, ds.label_names_to_use));
  259. % if a label_names_to_use name contains a string that is not one of the strings in the_labels, print an error message
  260. inds_of_bad_string_to_use_names = find(~ismember(ds.label_names_to_use, string_to_number_mapping));
  261. if ~isempty(inds_of_bad_string_to_use_names)
  262. bad_string_names = '';
  263. for iBadStringName = 1:length(inds_of_bad_string_to_use_names)
  264. bad_string_names = [bad_string_names ' ' ds.label_names_to_use{inds_of_bad_string_to_use_names(iBadStringName)} ','];
  265. end
  266. valid_string_names = '';
  267. for iValidString = 1:length(string_to_number_mapping)
  268. valid_string_names = [valid_string_names ' ' string_to_number_mapping{iValidString} ','];
  269. end
  270. error(['ds.label_string_names_to_use must be set to names in this list:' valid_string_names(1:end-1) '. ' ...
  271. 'The following ds.label_string_names_to_use strings not in the list:' bad_string_names(1:end-1)]);
  272. end
  273. ds.label_names_to_use = label_numbers_used; % convert given label names that should be used into numbers
  274. else
  275. ds.label_names_to_use = 1:length(string_to_number_mapping);
  276. end
  277. end
  278. % if ds.randomly_shuffle_labels_before_running == 1, randomly shuffle the labels the first time get_data is run
  279. if (ds.randomly_shuffle_labels_before_running == 1)
  280. 'randomly shuffling the labels'
  281. % added in NDT version 1.0.2 so that for simultaneously recorded populations the labels in all sites are shuffled the same way (previous version returned an error when shuffling labels on simultaneously recorded populations)
  282. if ds.create_simultaneously_recorded_populations > 0
  283. shuffled_labels = ds.the_labels{1}(randperm(length(ds.the_labels{1}))); % all sites should have the same labels, so will shuffle the labels for the first site only and will use this order for all sites
  284. for iSite = 1:length(ds.the_labels)
  285. ds.the_labels{iSite} = shuffled_labels;
  286. end
  287. % for non-simultaneously recorded datasets, shuffle each channel separately (same as NDT version 1.0.0)
  288. else
  289. for iSite = 1:length(ds.the_labels) % will shuffle the labels from all sites, not just from those specified in sites_to_use
  290. ds.the_labels{iSite} = ds.the_labels{iSite}(randperm(length(ds.the_labels{iSite})));
  291. end
  292. end
  293. end
  294. % if using simultaneously recorded populations, convert data to a format that will make code run a little faster
  295. if ds.create_simultaneously_recorded_populations == 0
  296. if ~(iscell(ds.the_data))
  297. the_data = ds.the_data;
  298. the_labels = ds.the_labels;
  299. for iSite = 1:size(the_data, 2)
  300. curr_data{iSite} = squeeze(the_data(:, iSite, :));
  301. curr_labels{iSite} = ds.the_labels;
  302. end
  303. ds.the_data = curr_data;
  304. ds.the_labels = curr_labels;
  305. end
  306. elseif ds.create_simultaneously_recorded_populations > 0
  307. if iscell(ds.the_data)
  308. simultaneous_labels_to_use = ds.the_labels{1}; % this should be ok, since all channels should have all the labels (regardless of whether a channel will ultimately be used)
  309. the_data = ds.the_data;
  310. for iSite = 1:length(the_data)
  311. if sum(abs(simultaneous_labels_to_use - ds.the_labels{iSite})) ~= 0
  312. error('problem, all simultaneously recorded neurons should have the same labels')
  313. end
  314. the_simultaneous_data(:, :, iSite) = the_data{iSite};
  315. end
  316. ds.the_data = permute(the_simultaneous_data, [1 3 2]);
  317. ds.the_labels = simultaneous_labels_to_use;
  318. end
  319. end
  320. if isempty(ds.time_periods_to_get_data_from)
  321. if iscell(ds.the_data) % for pseudo-populations
  322. num_time_periods = size(ds.the_data{1}, 2);
  323. else % for simultaneous data
  324. num_time_periods = size(ds.the_data, 3);
  325. end
  326. for i = 1:num_time_periods
  327. time_periods_to_use{i} = i;
  328. end
  329. ds.time_periods_to_get_data_from = time_periods_to_use;
  330. end
  331. % now that everything has been initialized, set inialized flag to 1
  332. ds.initialized = 1;
  333. end % end initialization
  334. % access to objects' fields in matlab is very slow (which is super pathetic), so to get the code to run faster I have use temporary copies of the data
  335. % (hopefully matlab will fix this in the future)
  336. the_data = ds.the_data;
  337. the_labels = ds.the_labels;
  338. num_cv_splits = ds.num_cv_splits;
  339. num_times_to_repeat_each_label_per_cv_split = ds.num_times_to_repeat_each_label_per_cv_split;
  340. curr_resample_sites_to_use = ds.curr_resample_sites_to_use;
  341. label_names_to_use = ds.label_names_to_use; if size(label_names_to_use, 1) ~= 1, label_names_to_use = label_names_to_use'; end % make sure labels numbers are in the correct orientation
  342. sites_to_use = ds.sites_to_use;
  343. sites_to_exclude = ds.sites_to_exclude;
  344. num_resample_sites = ds.num_resample_sites;
  345. sample_sites_with_replacement = ds.sample_sites_with_replacement;
  346. create_simultaneously_recorded_populations = ds.create_simultaneously_recorded_populations;
  347. %use_random_subset_of_k_labels_each_time_data_is_retrieved = ds.use_random_subset_of_k_labels_each_time_data_is_retrieved;
  348. % a santy checks
  349. if isempty(sites_to_use)
  350. error('sites_to_use can not be empty')
  351. end
  352. % if sites_to_use is a number that is less than 0, use all sites
  353. if (sites_to_use < 1)
  354. if create_simultaneously_recorded_populations > 0
  355. sites_to_use = 1:size(the_data, 2);
  356. else
  357. sites_to_use = 1:length(the_data);
  358. end
  359. end
  360. if ~isempty(sites_to_exclude) % can exclude specific neurons as well as specify which ones should be used
  361. sites_to_use = setdiff(sites_to_use, sites_to_exclude);
  362. end
  363. if isempty(label_names_to_use) || (isscalar(label_names_to_use) && label_names_to_use < 1)
  364. if create_simultaneously_recorded_populations > 0
  365. label_names_to_use = unique(the_labels);
  366. else
  367. label_names_to_use = unique(the_labels{sites_to_use(1)}); % use all the labels as the default value (assuming that the first used neuron 1 has all the labels shown - which might not be a foolproof assumption) % changed on 1/25/12
  368. end
  369. end
  370. % more sanity checks
  371. if length(label_names_to_use) ~= length(unique(label_names_to_use))
  372. warning('some labels were listed twice in the field ds.label_names_to_use, (these duplicate enteries will be ignored)');
  373. label_names_to_use = unique(label_names_to_use);
  374. end
  375. % making sure create_simultaneously_recorded_populations is a valid argument
  376. if (create_simultaneously_recorded_populations > 2) || (create_simultaneously_recorded_populations < 0)
  377. error('create_simultaneously_recorded_populations must be set to 0, 1 or 2');
  378. end
  379. % if the number of resample neurons is not specified, use all neurons
  380. if num_resample_sites < 1
  381. num_resample_sites = length(sites_to_use);
  382. end
  383. % code for randomly selecting k labels to use each time data is retrieved
  384. %if use_random_subset_of_k_labels_each_time_data_is_retrieved > 1 % needs at least 2 labels for a classification problem to work
  385. % rand_label = label_names_to_use(randperm(length(label_names_to_use)));
  386. % label_names_to_use = rand_label(1:use_random_subset_of_k_labels_each_time_data_is_retrieved);
  387. %end
  388. % if specific sites to be used have not been given (as should usually be the case), randomly select some sites to use from the larger population
  389. if length(curr_resample_sites_to_use) == 1 && (curr_resample_sites_to_use < 1) %isempty(curr_resample_sites_to_use)
  390. if ~(sample_sites_with_replacement) % only use each feature once in a population vector
  391. curr_resample_sites_to_use = sites_to_use(randperm(length(sites_to_use)));
  392. curr_resample_sites_to_use = sort(curr_resample_sites_to_use(1:num_resample_sites)); % sorting just for the heck of it
  393. else % selecting random features with replacement (i.e., the same feature can be repeated multiple times in a population vector).
  394. initial_inds = ceil(rand(1, num_resample_sites) * num_resample_sites); % can have multiple copies of the same feature within a population vector
  395. curr_resample_sites_to_use = sort(sites_to_use(initial_inds));
  396. end
  397. end
  398. % making code more robust in case transpose of curr_resample_sites_to_use is actually passed as an argument
  399. if (size(curr_resample_sites_to_use, 1) > 1)
  400. curr_resample_sites_to_use = curr_resample_sites_to_use';
  401. end
  402. % pre-allocate memory
  403. all_data_point_labels = NaN .* ones(length(unique(label_names_to_use)) * num_cv_splits * num_times_to_repeat_each_label_per_cv_split, 1);
  404. start_boostrap_ind = 1;
  405. % make sure label_names_to_use is a row vector
  406. if size(label_names_to_use, 1) > 1
  407. label_names_to_use = label_names_to_use';
  408. end
  409. if create_simultaneously_recorded_populations == 0
  410. % pre-allocate memory.
  411. the_resample_data = NaN .* ones(length(unique(label_names_to_use)) * num_cv_splits * num_times_to_repeat_each_label_per_cv_split, length(curr_resample_sites_to_use), size(the_data{1}, 2));
  412. % if someone has changed the data or the labels after they have already set create_simultaneously_recorded_populations = 0, then the format of these variables needs to be converted back
  413. if ~(iscell(ds.the_data)) || ~(iscell(the_labels))
  414. create_simultaneously_recorded_populations = 0;
  415. end
  416. % create a 3 dimensional tensor the_resample_data that is [(num_labels * num_cv_slits * num_repeats_per_cv_label) x num_neurons x num_time_bins] large
  417. for iLabel = label_names_to_use
  418. cNeuron = 1;
  419. for iNeuron = unique(curr_resample_sites_to_use)
  420. % choose random trials to use for each label type
  421. curr_trials_to_use = find(the_labels{iNeuron} == iLabel); %find(ds.the_labels{iNeuron} == iLabel);
  422. curr_trials_to_use = curr_trials_to_use(randperm(length(curr_trials_to_use)));
  423. if length(curr_trials_to_use) < (num_cv_splits * num_times_to_repeat_each_label_per_cv_split)
  424. error(['Requestion data from more trials of a given condition than has been recorded. This is due to ' ...
  425. '(ds.num_cv_splits * ds.num_times_to_repeat_each_label_per_cv_split) being greater than the number of times a given condition ' ...
  426. 'is present in the data (for at least one site). Make sure that only sites that have enough repetitions of each condition are used ' ...
  427. '(this can be done by setting ds.site_to_use = find_sites_with_at_least_k_repeats_of_each_label(the_labels_to_use, num_cv_splits) )']);
  428. return;
  429. else
  430. curr_trials_to_use = curr_trials_to_use(1:(num_cv_splits * num_times_to_repeat_each_label_per_cv_split));
  431. end
  432. % put everything into the correct number of CV splits
  433. for iRepeats = 1:length(find(curr_resample_sites_to_use == iNeuron))
  434. the_resample_data(start_boostrap_ind:(start_boostrap_ind + length(curr_trials_to_use) - 1), cNeuron, :) = the_data{iNeuron}(curr_trials_to_use, :);
  435. cNeuron = cNeuron + 1;
  436. end
  437. end % end iNeuron
  438. all_data_point_labels(start_boostrap_ind:(start_boostrap_ind + length(curr_trials_to_use) - 1)) = iLabel .* ones(length(start_boostrap_ind:(start_boostrap_ind + length(curr_trials_to_use) - 1)), 1);
  439. start_boostrap_ind = start_boostrap_ind + length(curr_trials_to_use);
  440. end % end for iLabel
  441. % if creating simultaneously recorded populations ...
  442. elseif create_simultaneously_recorded_populations > 0
  443. the_resample_data = NaN .* ones(length(unique(label_names_to_use)) * num_cv_splits * num_times_to_repeat_each_label_per_cv_split, length(curr_resample_sites_to_use), size(the_data, 3));
  444. the_data = the_data(:, curr_resample_sites_to_use, :);
  445. % choose (num_cv * num_repeats) random data points for each class
  446. for iLabel = label_names_to_use
  447. % choose random trials to use for each label type
  448. curr_trials_to_use = find(the_labels == iLabel);
  449. curr_trials_to_use = curr_trials_to_use(randperm(length(curr_trials_to_use)));
  450. if length(curr_trials_to_use) < (num_cv_splits * num_times_to_repeat_each_label_per_cv_split)
  451. error('problems: asking for more trials of a given stimuli type then were recorded in the experiment'); % maybe this should be an error
  452. return;
  453. else
  454. curr_trials_to_use = curr_trials_to_use(1:(num_cv_splits * num_times_to_repeat_each_label_per_cv_split));
  455. end
  456. the_resample_data(start_boostrap_ind:(start_boostrap_ind + length(curr_trials_to_use) - 1), :, :) = the_data(curr_trials_to_use, :, :);
  457. all_data_point_labels(start_boostrap_ind:(start_boostrap_ind + length(curr_trials_to_use) - 1)) = iLabel .* ones(length(start_boostrap_ind:(start_boostrap_ind + length(curr_trials_to_use) - 1)), 1);
  458. start_boostrap_ind = start_boostrap_ind + length(curr_trials_to_use);
  459. end
  460. end % end simultaneous populations
  461. clear the_data % clear up some memory
  462. all_resample_data_inds = 1:size(the_resample_data, 1);
  463. time_periods_to_get_data_from = ds.time_periods_to_get_data_from;
  464. % convert the_resample_data into a training and splits
  465. for iTimePeriod = 1:length(time_periods_to_get_data_from)
  466. curr_data = the_resample_data(:, :, time_periods_to_get_data_from{iTimePeriod});
  467. the_resample_data_time = reshape(curr_data, [size(curr_data, 1) size(curr_data, 2) * size(curr_data, 3)]);
  468. if (create_simultaneously_recorded_populations > 1)
  469. the_site_ids = 1:size(curr_data, 2);
  470. simul_to_pseudo_feature_to_siteID_mapping{iTimePeriod} = repmat(the_site_ids, [1 size(curr_data, 3)]);
  471. end
  472. cv_start_ind = 1;
  473. for iCV = 1:num_cv_splits
  474. curr_cv_inds = []; %NaN .* ones(num_times_to_repeat_each_label_per_cv_split * length(ds.label_names_to_use), 1);
  475. for iNumRepeatsPerLabel = 1:num_times_to_repeat_each_label_per_cv_split
  476. curr_cv_inds = [curr_cv_inds cv_start_ind:(num_cv_splits * num_times_to_repeat_each_label_per_cv_split):size(the_resample_data, 1)];
  477. cv_start_ind = cv_start_ind + 1;
  478. end
  479. % these cells arrays contain the data for each CV splits separately,
  480. % but don't need to create these, rather this function will just return the CV data divided into training and test sets
  481. % % % cross_validation_splits_all_time_periods{iCV}{iTimePeriod} = the_resample_data_time(curr_cv_inds, :); % old get_resample_data7 format...
  482. % cross_validation_splits_all_time_periods{iTimePeriod}{iCV} = the_resample_data_time(curr_cv_inds, :)'; % should replace above with this soon
  483. % cross_validation_labels{iCV} = all_data_point_labels(curr_cv_inds); % not sure why I need a separate one for each CV?
  484. % can actually just return these instead...
  485. XTr_all_time_cv{iTimePeriod}{iCV} = the_resample_data_time(setdiff(all_resample_data_inds, curr_cv_inds), :)';
  486. XTe_all_time_cv{iTimePeriod}{iCV} = the_resample_data_time(curr_cv_inds, :)';
  487. %YTr_all{iTimePeriod} = all_data_point_labels(setdiff(all_resample_data_inds, curr_cv_inds));
  488. %YTe_all{iTimePeriod} = all_data_point_labels(curr_cv_inds);
  489. % might as well return these as vectors (rather than cell arrays) since they are the same at all time periods
  490. YTr_all = all_data_point_labels(setdiff(all_resample_data_inds, curr_cv_inds));
  491. YTe_all = all_data_point_labels(curr_cv_inds);
  492. end
  493. end
  494. % If create_simultaneously_recorded_populations == 2, create pseudo-populations for training, and simultaneous data for testing.
  495. % This is useful for assessing I_diag as described by Averbeck, Latham and Pouget, Nature Neurosience, May 2006.
  496. if create_simultaneously_recorded_populations == 2
  497. XTr_all_time_cv = turn_training_simultaneous_data_into_pseudo_populations(XTr_all_time_cv, YTr_all, simul_to_pseudo_feature_to_siteID_mapping);
  498. end
  499. % This is another type of sanity check on pseudo-populations, but really is no reason to use this, so I am not going to give this as an option for now
  500. % if create_simultaneously_recorded_populations == 3
  501. % [XTr_all_time_cv XTe_all_time_cv] = turn_all_simultaneous_data_into_pseudo_populations(XTr_all_time_cv, YTr_all, XTe_all_time_cv, YTe_all, simul_to_pseudo_feature_to_siteID_mapping)
  502. % end
  503. end % end get_data
  504. end % end methods
  505. end % end class

basic_DS.m, no license · at the source

Overview

  1. National Institute on Drug Abuse, Intramural Research Program Baltimore United States
  2. State Key Laboratory of Cognitive Neuroscience and Learning, Beijing Normal University and Chinese Institute for Brain Research Beijing China
Institutions: National Institute on Drug Abuse (United States); Beijing Normal University (China)
Journal: eLife, volume 15, article RP109883
Dates: published online 21 April 2026
Type: Research article · Language: English
License: CC0
Identifiers: DOI 10.7554/elife.109883 · PMID 42011049 · PMCID PMC13099136 · OpenAlex W7123435098
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: extracellular electrophysiology (units, LFP) (modality), rat (organism), other condition (population), systems (subfield)
Methods: Spectral & time-frequency, Machine learning, Statistics, Preprocessing, Single-unit activity, calcium imaging
Keywords: cocaine, orbitofrontal, single unit, Rat
MeSH: Cocaine*, Neurons*, Prefrontal Cortex*, Animals, Male, Rats, Rats, Long-Evans, Self Administration (* major topic)
Journal subjects: Cell Biology
Topic: Memory and Neural Mechanisms (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Funding: National Institute on Drug Abuse (Z1A-DA000587)
Citations: cited by 1 paper (Europe PMC); 41 references in the paper

Abstract

The orbitofrontal cortex (OFC) is critical to identifying task structure and to generalizing appropriately across task states with similar underlying or hidden causes. This capability is at the heart of OFCs proposed role in a network responsible for cognitive mapping, and its loss can explain many deficits associated with OFC damage or inactivation. Substance use disorder is defined by behaviors that share much in common with these deficits, such as an inability to modify learned behaviors in the face of new information about undesired consequences. One explanation for this similarity would be if addictive drugs impacted the ability of OFC to recognize underlying similarities, hidden states, that allow information learned in one setting to be used in another. To explore this possibility, we trained rats to self-administer cocaine and then recorded single-unit activity in lateral OFC as these rats performed in an odor sequence task consisting of unique and shared positions. In well-trained controls, we observed chance decoding of sequence at shared positions and near chance decoding even at unique positions, reflecting the irrelevance of distinguishing these positions in the task. By contrast, in cocaine-experienced rats, decoding remained significantly elevated, particularly at the positions that had superficial sensory differences that were collapsed in controls across learning. These neural differences were accompanied by increases in behavioral variability at these positions. A tensor component analysis showed that this effect of reduced generalization after cocaine use also extended across positions in the sequences. These results show that prior cocaine use disrupts the normal identification of hidden states by OFC.

Reproduced under the paper's license (CC0), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.

OSF azvhm

License: none: the authors keep all their rights
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Size: 29 files
Software Heritage: not checked
Found in: “Data availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Statistics and Machine Learning Toolbox (27 files), fdr_bh (Benjamini-Hochberg FDR) (2 files), Matplotlib (1 file), NumPy (1 file), SciPy (1 file)
Availability: 1 check, the latest on 29 September 2026: the link answers (HTTP 200)
  • 29 September 2026: the link answers (HTTP 200)
216 files
At the source: osf.io/azvhm/

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 216 scripts, each with its path and the digest of its content;
  • 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability

Data and code availability: All data and analysis code associated with this study are available on OSF at https://osf.io/azvhm/.

The following dataset was generated:

ZongW Open Science Framework2026Prior cocaine use disrupts identification of hidden states by single units and neural ensembles in orbitofrontal cortexazvhm10.7554/eLife.109883PMC1309913642011049

Reproduced under the paper's license (CC0), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 29 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 5 authors, 4 keywords, 8 MeSH terms, 1 funder, 41 references.

Cite

This paper

Zong, W., Mueller, L., Zhang, Z., Zhou, J., & Schoenbaum, G. (2026). Prior cocaine use disrupts identification of hidden states by single units and neural ensembles in orbitofrontal cortex. eLife, 15, RP109883. https://doi.org/10.7554/elife.109883

BibTeX

@article{zong2026prior,
author = {Zong, Wenhui and Mueller, Lauren and Zhang, Zhewei and Zhou, Jinfeng and Schoenbaum, Geoffrey},
title = {{Prior cocaine use disrupts identification of hidden states by single units and neural ensembles in orbitofrontal cortex}},
journal = {eLife},
year = {2026},
month = apr,
volume = {15},
pages = {RP109883},
publisher = {eLife Sciences Publications, Ltd},
issn = {2050-084X},
doi = {10.7554/elife.109883},
url = {https://doi.org/10.7554/elife.109883},
pmid = {42011049},
pmcid = {PMC13099136}
}

RIS

TY - JOUR
AU - Zong, Wenhui
AU - Mueller, Lauren
AU - Zhang, Zhewei
AU - Zhou, Jinfeng
AU - Schoenbaum, Geoffrey
TI - Prior cocaine use disrupts identification of hidden states by single units and neural ensembles in orbitofrontal cortex
T2 - eLife
J2 - Elife
PY - 2026
DA - 2026/04/21
VL - 15
SP - RP109883
SN - 2050-084X
PB - eLife Sciences Publications, Ltd
DO - 10.7554/elife.109883
UR - https://doi.org/10.7554/elife.109883
LA - en
ER -

CSL-JSON

{
"id": "10.7554/elife.109883",
"type": "article-journal",
"title": "Prior cocaine use disrupts identification of hidden states by single units and neural ensembles in orbitofrontal cortex",
"container-title": "eLife",
"author": [
{
"family": "Zong",
"given": "Wenhui"
},
{
"family": "Mueller",
"given": "Lauren"
},
{
"family": "Zhang",
"given": "Zhewei"
},
{
"family": "Zhou",
"given": "Jinfeng"
},
{
"family": "Schoenbaum",
"given": "Geoffrey"
}
],
"container-title-short": "Elife",
"volume": "15",
"page": "RP109883",
"DOI": "10.7554/elife.109883",
"PMID": "42011049",
"PMCID": "PMC13099136",
"ISSN": "2050-084X",
"publisher": "eLife Sciences Publications, Ltd",
"URL": "https://doi.org/10.7554/elife.109883",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
21
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.cub.2026.05.068 [code]
An abstract relational map emerges in the human medial prefrontal cortex with consolidation.
Journal: Current biology : CB
In common: Statistics and Machine Learning Toolbox, SciPy, Matplotlib, 1 other tool, systems, 3 references
[2] doi:10.7554/elife.108223 [code]
Two time scales of adaptation in human learning rates.
Journal: eLife
In common: SciPy, Matplotlib, NumPy, 3 references
[3] doi:10.1371/journal.pbio.3003824 [code]
Flexible goal learning involves coordinated population activity in dCA1 and medial orbitofrontal cortex.
Journal: PLoS biology
In common: SciPy, Matplotlib, NumPy, rat, systems, 2 references
[4] doi:10.1523/jneurosci.2001-25.2026 [code]
Dynamics of Dentate Gyrus Place Cells and Dentate Spikes during Spatial and Nonspatial Changes in Environments.
Journal: The Journal of neuroscience : the official journal of the Society for Neuroscience
In common: Statistics and Machine Learning Toolbox, SciPy, Matplotlib, 1 other tool, extracellular electrophysiology (units, LFP), rat, systems
[5] doi:10.1016/j.crmeth.2026.101308 [code]
Projection targeting with phototagging to study the structure and function of retinal ganglion cells.
Journal: Cell reports methods
In common: Statistics and Machine Learning Toolbox, SciPy, Matplotlib, 1 other tool, extracellular electrophysiology (units, LFP), rat, systems
[6] doi:10.1038/s41467-026-77318-1 [code]
Offline generative network reconfiguration guides insight-like accelerated learning by assimilation into schema in rats.
Journal: Nature communications
In common: Statistics and Machine Learning Toolbox, SciPy, Matplotlib, rat, systems, 1 reference
[7] doi:10.1186/s12916-026-04903-y [code]
Structural connectome architecture and biological vulnerability shape cortical atrophy in cocaine use disorder.
Journal: BMC medicine
In common: Statistics and Machine Learning Toolbox, SciPy, Matplotlib, 1 other tool, other condition, 1 reference
[8] doi:10.1038/s41593-026-02359-0 [code]
The cross-site reproducibility of MRI morphometric phenotypes in psychiatric disorders.
Journal: Nature neuroscience
In common: fdr_bh (Benjamini-Hochberg FDR), Statistics and Machine Learning Toolbox, SciPy, 2 other tools
[9] doi:10.1038/s41467-026-75959-w [code]
Charting higher-order models of brain function beyond pairwise interactions.
Journal: Nature communications
In common: fdr_bh (Benjamini-Hochberg FDR), Statistics and Machine Learning Toolbox, SciPy, 2 other tools
[10] doi:10.1038/s41593-026-02357-2 [code]
Experience reorganizes content-specific memory traces in macaques.
Journal: Nature neuroscience
In common: fdr_bh (Benjamini-Hochberg FDR), Statistics and Machine Learning Toolbox, SciPy, 2 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.