OSCR

Brain-inspired spatial intelligence for embodied agents.

Code ↔ Paper

16 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 16 matches · 2 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
  1. [1] § Methods › Experimental details › Simulator and datasets ↔ LLMAgent.py, lines 145–204 · score 0.71 · extrinsic attributes, intrinsic attributes, surrounding environmental, language, scenes
  2. [2] § Methods › The BSC-Nav framework › Working memory ↔ LLMAgent.py, lines 208–270 · score 0.70 · kitchen island, confidence scores, target locations, prompts, GPT, matching
  3. [3] § Results › Universal navigation across modalities and granularities ↔ demo.py, lines 25–147 · score 0.67 · navigation trajectories, mp3d, hm3d, navigation tasks, benchmark, episodes
  4. [4] § Methods › Experimental details › Simulator and datasets ↔ textnav_benchmark.py, lines 89–156 · score 0.65 · extrinsic attributes, intrinsic attributes, hm3d, episodes, navigation, scenes
  5. [5] § Methods › Experimental details › Simulation environments ↔ demo.py, lines 25–147 · score 0.63 · Stable Diffusion, success distance, working memory, Medium, exploration, benchmark
  6. [6] § Methods › The BSC-Nav framework › Cognitive map ↔ GES_vlnce/memory.py, lines 834–895 · score 0.62 · radial distance, patch tokens, weight, height, depth, RGB
  7. [7] § Methods › The BSC-Nav framework › Cognitive map ↔ memory_2.py, lines 842–903 · score 0.62 · radial distance, patch tokens, weight, height, depth, RGB
  8. [8] § Results › Construction and exploitation of structured spatial memory in embodied agents ↔ LLMAgent.py, lines 208–270 · score 0.62 · confidence score, highest confidence, target locations, answering, semantic, memory
  9. [9] § Methods › Experimental details › Simulation environments ↔ args.py, the whole file · a weak match · score 0.60 · Stable Diffusion, success distance, working memory, Medium, benchmark, depth
  10. [10] § Methods › Experimental details › Real-world deployment ↔ habitat-lab/habitat/datasets/rearrange/navmesh_utils.py, lines 592–640 · score 0.58 · angular speed, linear speed, threshold, robot, configured, position
  11. [11] § Methods › Experimental details › Foundation model specifications ↔ GES_vlnce/memory.py, lines 39–144 · score 0.56 · YOLO World, WorldV2, quantized, patch, Diffusion, tokens
  12. [12] § Methods › Experimental details › Foundation model specifications ↔ memory_2.py, lines 39–145 · score 0.56 · YOLO World, WorldV2, quantized, patch, Diffusion, tokens
  13. [13] § Methods › The BSC-Nav framework › Low-level navigation policy generation ↔ src/esp/nav/GreedyFollower.h, lines 17–163 · score 0.56 · avoid obstacles, greedy, planner, shortest, sequence, motion
  14. [14] § Methods › Experimental details › Simulator and datasets ↔ args.py, the whole file · a weak match · score 0.55 · HM3D scenes, MP3D scenes, benchmarks, episodes, navigation
  15. [15] § Methods › The BSC-Nav framework › Working memory ↔ LLMAgent.py, lines 9–67 · score 0.54 · enriched descriptions, refine, texture, GPT, model
  16. [16] § Methods › Experimental details › Baseline methods and comparison ↔ GES_vlnce/VLN_CE/vlnce_baselines/models/seq2seq_policy.py, lines 52–179 · score 0.51 · action embeddings, networks, vision, encoding, encoders, sequences

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 991 lines · 47 KB · no license · 4 matches

  1. from openai import OpenAI
  2. import json
  3. from io import BytesIO
  4. import base64
  5. from PIL import Image
  6. # from qwen_vl_utils import process_vision_info
  7. def imagenary_helper_visaug(client, text_prompt, vis):
  8. base64_images = image_to_base64(vis)
  9. completion = client.chat.completions.create(
  10. model="gpt-4o",
  11. timeout=500,
  12. messages=[
  13. {"role": "system", "content": "You are a helpful assistant."},
  14. {"role": "user", "content": [
  15. {"type": "text", "text": """
  16. You are an expert in generating prompts for text-to-image models.
  17. Your task is to enhance a given original goal description, which often only mentions a general category or a simple phrase describing the target object, by incorporating detailed context from four current scene images. You will create a more imaginative and contextually enriched description, using elements observed in the scene to create a coherent and vivid visual. This description will be used to guide a text-to-image model to generate an image that aligns with the style and context of the current scene.
  18. It is crucial that the target object remains the primary visual focus of the image. To complete this task effectively, it's suggested to follow these steps:
  19. 1. **Understand the Environment**: Extract and comprehend details from the provided observation images, such as the overall style, decoration, and elements of the scene. For example, is this a modern home, a classical residence, an exhibition hall, or an office? Consider the style in detail, and understand the context of the environment.
  20. 2. **Expand the Original Description **: Based on the scene analysis from step 1, enrich the original description with finer details. This may include materials, colors, textures, placement, and environmental elements surrounding the target object.
  21. 3. **Maintain Visual Focus **: Ensure that any additional context or background details do not overshadow the main target object. The primary subject should remain the focal point of the generated image, using language that emphasizes its prominence in the scene.
  22. **Guidelines for Creating Enhanced Descriptions: **
  23. 1. Details: Include sensory details like colors, textures, lighting, and reflections.
  24. 2. Background Elements: Add appropriate background elements that complement the scene without detracting from the focus of the image.
  25. Focus Phrasing: Use language that naturally draws attention to the target object (e.g., "centered," "prominently placed," "as the focal point").
  26. 3. Balance: Strike a balance between richness and simplicity. The target object or scene should always dominate the final image.
  27. **Enhanced Description Output Requirements: **
  28. 1. Provide a refined and detailed description in English.
  29. 2. Ensure the enhancement creates a vivid, coherent, and engaging scene, supporting the original description.
  30. 3. Avoid overly complex narratives or elements that distract from the primary object or scene.
  31. 4. Keep the description concise, limiting it to 70 words or less.
  32. **Examples:**
  33. 1.
  34. Original Goal Description: A green vase.
  35. Enhanced Description: A vibrant green ceramic vase, with a glossy, smooth surface, placed centrally on a polished wooden table. Soft natural light illuminates the vase from the large window behind it, casting gentle shadows on the table. The surrounding room is decorated in minimalist modern style with neutral tones, ensuring the vase is the central focal point of the scene.
  36. 2.
  37. Original Goal Description: A armchair.
  38. Enhanced Description: A sleek, modern blue armchair, upholstered in soft velvet, positioned prominently in a stylish living room. The chair is placed near a large floor-to-ceiling window, allowing natural light to highlight its deep blue hue. The room features minimalist decor with white walls, light wood flooring, and a few abstract art pieces on the walls. The armchair stands out as the main focal point, inviting comfort and relaxation.
  39. 3.
  40. Original Goal Description: A desk.
  41. Enhanced Description: A robust metal desk, with a weathered, matte surface, placed against a brick wall in an industrial-style office. The desk features exposed steel legs, and on its surface lies a sleek laptop, a coffee mug, and a few scattered papers. Overhead, a vintage filament bulb hangs from a chain, casting a warm glow over the scene. The surrounding decor includes minimalistic shelves and a large plant in the corner, yet the desk remains the focal point in the room’s urban, raw atmosphere.
  42. """
  43. },
  44. {"type": "text", "text": f"""
  45. Now, the Original Goal Description is "{text_prompt}", and the observation images are:
  46. """
  47. },
  48. {"type": "text", "text": f"observation1:"},
  49. {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{base64_images[0]}"}},
  50. {"type": "text", "text": f"observation2:"},
  51. {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{base64_images[2]}"}},
  52. {"type": "text", "text": f"""
  53. please follow the above requirements and examples to enhance this description, think step by step and give your analysis process and the final enhancement description following the format:
  54. **analysis process**: [your analysis process here]
  55. **enhancement description**: [your enhancement description here]
  56. """
  57. }]
  58. }]
  59. )
  60. return completion.choices[0].message.content
  61. def imagenary_helper(client, text_prompt):
  62. completion = client.chat.completions.create(
  63. model="gpt-4o",
  64. timeout=500,
  65. messages=[
  66. {"role": "system", "content": "You are a helpful assistant."},
  67. {"role": "user", "content": f"""
  68. You are an expert in refining and elaborating simple scene or object descriptions for use in text-to-image generation.
  69. Your task is to take a given raw navigation target description—often just a short phrase mentioning an object or a simple scene—and enhance it by adding imaginative and contextual details.
  70. It is crucial that the original object(s) remain the visual focal point. To achieve this:
  71. 1.**Expand on the Original Description**: Add details about the object(s), their materials, colors, textures, and immediate surroundings.
  72. 2.**Maintain Visual Dominance**: Ensure that any additional context or setting you provide does not overshadow the main elements described in the original text.
  73. The main subjects should remain the most prominent features of the resulting image.
  74. 3.**Emphasize the Main Objects**: Use language that clearly highlights the original objects, ensuring they stand out as the central focus.
  75. **Output Requirements:**
  76. 1.Provide a refined, detailed description in English.
  77. 2.Ensure the enhancements create a vivid, coherent, and inviting scene that supports the original description.
  78. 3.Avoid overly intricate storytelling or elements that distract from the original object or scene.
  79. 4.Please avoid overly long descriptions, limit descriptions to 70 words or less.
  80. **Guidelines for Creating Enhanced Descriptions:**
  81. 1.**Detailing**: Incorporate sensory details such as color, texture, lighting, and reflections.
  82. 2.**Contextual Elements**: Add subtle background elements or environmental details that complement but do not compete with the main subject.
  83. 3.**Focus Phrasing**: Use language that naturally draws attention to the original object(s) (e.g., “centered,” “prominently placed,” “serving as the focal point”).
  84. 4.**Balance**: Maintain a balance between enrichment and simplicity. The primary object or scene described in the original prompt should always dominate the final image.
  85. """
  86. },
  87. {"role": "user", "content": f"""
  88. **Examples**:
  89. 1.
  90. Original Description:
  91. a TV screen above cabinets.
  92. Enhanced Description:
  93. A sleek, flat-screen TV mounted above a set of smooth, white cabinetry. The TV's reflective surface catches the soft glow of recessed lighting,
  94. while the clean, minimalist cabinets provide a neat base that keeps the television as the primary visual anchor.
  95. 2.
  96. Original Description:
  97. a marble island in kitchen.
  98. Enhanced Description:
  99. A polished white marble island centered in a modern kitchen. Delicate gray veining runs across its surface, subtly reflecting under the warm overhead lights.
  100. Simple barstools and neat, neutral-toned countertops frame the island, ensuring it remains the kitchen's focal point.
  101. 3.
  102. Original Description:
  103. a coffee mug on a desk.
  104. Enhanced Description:
  105. A sturdy ceramic coffee mug, ivory in color, resting on a clean, wooden desk. Soft light from a nearby window gently highlights the mug's curved handle and the steam rising from its freshly poured contents.
  106. A simple laptop and a neatly arranged notepad stay in the background, making the mug stand out as the central feature.
  107. 4.
  108. Original Description:
  109. lamp.
  110. Enhanced Description:
  111. A classic white table lamp with a slender stem and an elegant lampshade, standing on a minimalist nightstand. The soft glow from the lamp illuminates its graceful lines,
  112. while the uncluttered background ensures the lamp remains the central feature of the scene.
  113. 5.
  114. Original Description:
  115. a painting.
  116. Enhanced Description:
  117. A colorful, abstract painting mounted on a textured, red-brick wall. The painting's bold brushstrokes and vibrant hues stand out sharply against the wall's rough surface.
  118. Soft track lighting above gently illuminates the artwork, drawing the eye directly to it.
  119. """
  120. },
  121. {"role": "user", "content": f"""
  122. Now, the original description is "{text_prompt}", please follow the above requirements and examples to enhance this description, and directly output the enhanced description.
  123. """
  124. },
  125. ])
  126. return completion.choices[0].message.content
  127. def imagenary_helper_long_text(client, text_prompt):
  128. goal_text_intrinsic, goal_text_extrinsic = text_prompt[0], text_prompt[1]
  129. completion = client.chat.completions.create(
  130. model="gpt-4o",
  131. timeout=500,
  132. messages=[
  133. {"role": "system", "content": "You are a helpful assistant."},
  134. {"role": "user", "content": f"""
  135. You are an expert in refining and elaborating simple scene or object descriptions for use in text-to-image generation.
  136. Your task is to reasonably merge the text description of the object instance and the text description of the surrounding environment, and output a complete description text to guide a text-to-image model to generate a photo of the object. It is crucial that the instance object(s) remain the visual focal point. To achieve these, you should:
  137. 1.**Expand on the Original Description**: Based on the original object description and environment description, you should provides a more fine-grained description of materials, colors, and textures.
  138. 2.**Maintain Visual Dominance**: Ensure that any additional context or setting you provide does not overshadow the main elements described in the original text. the main object should remain the most prominent features of the resulting image.
  139. **Output Requirements:**
  140. 1.Provide a refined, detailed description in English.
  141. 2.Ensure the enhancements create a vivid, coherent, and inviting scene that supports the original description.
  142. 3.Avoid overly intricate storytelling or elements that distract from the original object or scene.
  143. 4.The extended description must follow the original description and not contradict it, such as color, material, and appearance.
  144. 4.Please avoid overly long descriptions, limit descriptions to 70 words or less.
  145. **Guidelines for Creating Enhanced Descriptions:**
  146. 1.**Detailing**: Incorporate sensory details such as color, texture, lighting, and reflections.
  147. 2.**Contextual Elements**: Add subtle background elements or environmental details that complement but do not compete with the main subject.
  148. 3.**Focus Phrasing**: Use language that naturally draws attention to the original object(s) (e.g., “centered,” “prominently placed,” “serving as the focal point”).
  149. 4.**Balance**: Maintain a balance between enrichment and simplicity. The primary object or scene described in the original prompt should always dominate the final image.
  150. """
  151. },
  152. {"role": "user", "content": f"""
  153. **Examples**:
  154. "intrinsic_attributes": "According to the image description, this chair is an old wooden structure with blue seat cushions. Its specific material and color cannot be determined because there are no other details provided about its body or frame."
  155. "extrinsic_attributes": "There are no other objects around the chair in this image. Only a very simple environment is shown, with only one object: an old wooden table and four blue chairs."
  156. Enhanced Description Output:
  157. A weathered wooden chair with a timeworn frame and soft blue cushions sits prominently at the center. The rich grain of the wood contrasts with the smooth fabric, which shows subtle signs of wear. The surroundings are minimal, featuring a matching wooden table and four identical blue chairs, creating a serene and uncomplicated environment where the chair remains the clear focal point. The warm, natural light highlights the texture of the wood and fabric.
  158. """
  159. },
  160. {"role": "user", "content": f"""
  161. Now, the "intrinsic_attributes" is "{goal_text_intrinsic}", the "extrinsic_attributes" is "{goal_text_extrinsic}", please follow the above requirements and examples to enhance this description, and directly output the enhanced description.
  162. """
  163. },
  164. ])
  165. return completion.choices[0].message.content
  166. def long_memory_localized(client, text_prompt, long_memory):
  167. completion = client.chat.completions.create(
  168. model="gpt-4o",
  169. timeout=500,
  170. messages=[
  171. {"role": "system", "content": "You are a helpful assistant."},
  172. {"role": "user", "content": """
  173. You are an LLM Agent with a specific goal: given a textual description of a navigation target (e.g., “A marble island in a kitchen.”)
  174. and a memory list containing instances of detected objects, each with a label, a 3D location (three numerical coordinates), and a confidence score,
  175. you need to determine the most suitable memory instance to fulfill the navigation request.
  176. **Your memory data structure:**
  177. The memory is a list of objects in the environment, where each object is represented as a JSON-like structure of the following form:
  178. {
  179. "label": "<string>",
  180. "loc": [<float or int>, <float or int>, <float or int>],
  181. "confidence": <float>
  182. }
  183. label: A textual label describing the object (e.g., “a tv”, “a kitchen island”).
  184. loc: A three-element array representing the coordinates of that object in the environment.
  185. confidence: A floating-point value indicating how confident the system is in identifying this object as described by label.
  186. **Your task:**
  187. 1.Understand the target description: You will be given a textual goal description of a navigation target, such as “A marble island in a kitchen.”
  188. Your first step is to interpret this description and deduce which object label from the memory best matches it semantically.
  189. For example, if the target is “A marble island in a kitchen,” and you have memory instances labeled “a kitchen island” or “a marble island,”
  190. you should identify that these instances correspond to the target description. Consider synonyms and close matches. If no exact label is found,
  191. choose the label that is most semantically similar to the target description. For instance, if the target mentions “a marble island” and the memory only has “a kitchen island,”
  192. you should still select the “a kitchen island” label as it is likely the intended object.
  193. 2.Identify the relevant instances: Once you have determined the best matching label, filter the memory list to only those instances whose label matches (or closely matches) that label.
  194. 3.Evaluate confidence and consolidate duplicates: Among these filtered instances, consider that multiple memory entries may actually represent the same object,
  195. possibly due to partial overlaps or multiple detections.
  196. - Look at their loc coordinates. If multiple instances with the same label have very close or nearly identical coordinates, treat them as the same object.
  197. - Determine which set of coordinates (if there are multiple distinct sets) is the most reliable representation of the object. Reliability is judged primarily by the highest confidence value. If multiple instances cluster together with similar locations, select the one with the highest confidence or, if confidence is similar, the one that best aligns with the object as described.
  198. 4.Select the final loc: After you have grouped instances and decided which group best represents the target object, output the coordinates (loc) of the best match.
  199. If multiple objects (>=3 items) match the description equally well, choose the three coordinates (loc) with the highest confidence.
  200. 5.Produce a final answer: Return the selected location coordinates as the final answer, (important!!) must be in the format '**Result**: (Nav Loc 1: [...], Nav Loc 2: [...], Nav Loc 3: [...])' or '**Result**: (Nav Loc: Unable to find)'.
  201. **Important details:**
  202. - Always provide reasoning internally (you may do it in hidden scratchpads if available) before giving the final result.
  203. The final user-visible answer should be concise and directly address the task.
  204. - If no objects are found that are semantically relevant to the target description, explicitly indicate that no suitable object was found.
  205. - Follow these steps for every input you receive.
  206. """
  207. },
  208. {"role": "user", "content": f"navigation target:{text_prompt}"},
  209. {"role": "user", "content": f"memory:{long_memory}"},
  210. {"role": "user", "content": "Now please start thinking one step at a time and then Briefly tell me the target location I need to go to and return as '**Result**: (Nav Loc 1: [...], Nav Loc 2: [...], Nav Loc 3: [...])' format, If there is no suitable target in memory, return as '**Result**: (Nav Loc: Unable to find)"},
  211. ])
  212. return completion.choices[0].message.content
  213. def image_to_base64(images: Image.Image, fmt="JPEG") -> str:
  214. base64_images = []
  215. for img in images:
  216. output_buffer = BytesIO()
  217. img.save(output_buffer, format=fmt)
  218. byte_data = output_buffer.getvalue()
  219. base64_str = base64.b64encode(byte_data).decode('utf-8')
  220. base64_images.append(base64_str)
  221. return base64_images
  222. # def succeed_determine(client, text_prompt, obs):
  223. # base64_image = image_to_base64(obs)
  224. # completion = client.chat.completions.create(
  225. # model="gpt-4o",
  226. # messages=[
  227. # {"role": "system", "content": "You are a helpful assistant."},
  228. # {"role": "user", "content": """
  229. # You are an AI assistant tasked with assessing whether a navigation task has been successfully completed. You will be given:
  230. # 1.A observation image (the image observed at the end of navigation).
  231. # 2.A text description of the target object the navigation aimed to locate.
  232. # Your goal is to determine:
  233. # 1.Explanation: Provide a concise explanation showing how you arrived at your conclusion, referencing any matching or non-matching details between the final observation and the target description.
  234. # 2.Success or not: Does the final observation indicate that the agent has found the target object?
  235. # **Important Requirements**
  236. # Before your success verdict, provide an explanation on a line, starting with Explanation: and then your reasoning.
  237. # If the final observation includes the target object based on the textual description, respond with exactly Success: yes.
  238. # If the final observation does not include the target object, respond with exactly Success: no.
  239. # Do not add extra lines or deviate from the required format.
  240. # **Format of Your Answer**
  241. # First line: Explanation: [Your explanation here]
  242. # Second line: Success: yes OR Success: no
  243. # """
  244. # },
  245. # {"role": "user",
  246. # "content": [
  247. # {
  248. # "type": "text",
  249. # "text": "observation image:"},
  250. # {
  251. # "type": "image_url",
  252. # "image_url": {"url": f"data:image/jpeg;base64,{base64_image}"},},
  253. # ]
  254. # },
  255. # {"role": "user", "content": f"target description:{text_prompt}"},
  256. # {"role": "user", "content": "Now please start thinking step by step"},
  257. # ])
  258. # return completion.choices[0].message.content
  259. def succeed_determine(client, text_prompt, obss):
  260. base64_images = image_to_base64(obss)
  261. messages = [
  262. {"role": "system", "content": "You are a helpful assistant."},
  263. {"role": "user", "content": [
  264. {
  265. "type": "text", "text":
  266. """
  267. You will be provided with 2 navigation observation images from from different perspectives of the agent, and a textual description of the navigation goal. Please follow the steps below to determine whether the current navigation task is successful:
  268. 1.Determine Target Presence: Analyze the provided images one by one, to ascertain whether the navigation goal is present in these images. This means evaluating if the agent has arrived near the target location.
  269. 2. Output Format:
  270. First Line: Success: yes OR Success: no
  271. Second Line: Give your analysis results in detail
  272. Examples
  273. '''
  274. Success: yes
  275. [analysis results]
  276. '''
  277. or
  278. '''
  279. Success: no
  280. [analysis results]
  281. '''
  282. Please analyze according to the above requirements and respond strictly in the specified format.
  283. """
  284. },
  285. {
  286. "type": "text", "text": f"target description: {text_prompt}"
  287. }
  288. ]
  289. },]
  290. for idx, base64_image in enumerate(base64_images):
  291. messages[1]["content"].extend([
  292. {"type": "text", "text": f"observation view {idx}:"},
  293. {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{base64_image}"}}
  294. ])
  295. messages[1]["content"].extend([
  296. {"type": "text", "text": "Now please start thinking step by step"},
  297. ])
  298. completion = client.chat.completions.create(
  299. model="gpt-4o",
  300. timeout=500,
  301. messages=messages
  302. )
  303. return completion.choices[0].message.content
  304. def succeed_determine_singleview(client, text_prompt, obss):
  305. base64_images = image_to_base64(obss)
  306. messages = [
  307. {"role": "system", "content": "You are a helpful assistant."},
  308. {"role": "user", "content": [
  309. {
  310. "type": "text", "text":
  311. """
  312. You will be provided with a navigation observation image from a robot, and a textual description of the navigation goal. Please follow the steps below to determine whether the current navigation task is successful:
  313. 1.Determine Target Presence: Analyze the provided images to ascertain whether the navigation goal is present in these images And close enough (within 2 meters). This means evaluating if the robot has arrived near the target location. Be careful not to misclassify the similar categories (e.g. sofas and chairs are easily confused)
  314. 2.Determine whether need to move forward. If you have found the target according to the step 1, you need to further determine whether you need to move forward a small step to get closer to the target object. If need to, answer 'need forward: yes'. If you think you are close enough (within 1m), answer 'need forward: no'.
  315. 3. Output Format:
  316. First Line: Success: yes OR Success: no
  317. Second Line (only when 'success: yes'): need forward: yes OR need forward: no
  318. Third Line: Give your analysis results in detail
  319. Examples
  320. '''
  321. Success: yes
  322. need forward: yes
  323. [analysis results]
  324. '''
  325. or
  326. '''
  327. Success: yes
  328. need forward: no
  329. [analysis results]
  330. '''
  331. or
  332. '''
  333. Success: no
  334. [analysis results]
  335. '''
  336. Please analyze according to the above requirements and respond strictly in the specified format.
  337. """
  338. },
  339. {
  340. "type": "text", "text": f"target description: {text_prompt}"
  341. }
  342. ]
  343. },]
  344. for idx, base64_image in enumerate(base64_images):
  345. messages[1]["content"].extend([
  346. {"type": "text", "text": f"observation view:"},
  347. {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{base64_image}"}}
  348. ])
  349. messages[1]["content"].extend([
  350. {"type": "text", "text": "Now please start thinking step by step"},
  351. ])
  352. completion = client.chat.completions.create(
  353. model="gpt-4o",
  354. timeout=500,
  355. messages=messages
  356. )
  357. return completion.choices[0].message.content
  358. def succeed_determine_singleview_with_imggoal(client, img_prompt, obss):
  359. obss = [obs.resize((512, 512)) for obs in obss]
  360. img_prompt = img_prompt.resize((512, 512))
  361. base64_images = image_to_base64(obss)
  362. base64_images_goal = image_to_base64([img_prompt])
  363. messages = [
  364. {"role": "system", "content": "You are a helpful assistant."},
  365. {"role": "user", "content": [
  366. {
  367. "type": "text", "text":
  368. """
  369. You will get an image of the navigation target and an image of the current observation from the robot.
  370. 1. Analyze the main instance objects contained in the two images, especially the closest objects. Think about what objects they are? What appearance features do they have?
  371. 2. Compare the two images, combined with the analysis of the step 1, to determine whether the current robot has reached the vicinity of the target. This means that the current observation image is taken at the location of the navigation target image.
  372. 3. Determine whether need to move forward. If you think the robot has reached the target, you need to further determine whether it needs to move forward a small step to get closer to the target object.
  373. Note: The viewpoints of the two images are usually different. You need to judge carefully to avoid misjudgment.
  374. Output Format:
  375. First Line: Success: yes OR Success: no
  376. Second Line (only when 'success: yes'): need forward: yes OR need forward: no
  377. Third Line: Give your analysis results in detail
  378. Examples
  379. '''
  380. Success: yes
  381. need forward: yes
  382. [analysis results]
  383. '''
  384. or
  385. '''
  386. Success: yes
  387. need forward: no
  388. [analysis results]
  389. '''
  390. or
  391. '''
  392. Success: no
  393. [analysis results]
  394. '''
  395. Please analyze according to the above requirements and respond strictly in the specified format.
  396. """
  397. },
  398. {
  399. "type": "text", "text": f"navigation goal image:"
  400. },
  401. {
  402. "type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{base64_images_goal[0]}"}
  403. },
  404. {
  405. "type": "text", "text": f"current observation image:"
  406. },
  407. {
  408. "type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{base64_images[0]}"}
  409. },
  410. {
  411. "type": "text", "text": "Now please start analysing."
  412. },
  413. ],
  414. }
  415. ]
  416. completion = client.chat.completions.create(
  417. model="gpt-4o",
  418. timeout=500,
  419. messages=messages
  420. )
  421. return completion.choices[0].message.content
  422. def touching_helper(client, text_prompt, obss, model=None, processor=None):
  423. base64_image = image_to_base64(obss)[0]
  424. messages = [
  425. {"role": "user",
  426. "content":
  427. [
  428. {
  429. "type": "text", "text":
  430. """
  431. Suppose you are an agent performing a navigation task, and you need to make yourself as close to a specified target as possible. Now that you have reached the vicinity of the target, I will provide you with a description of the target object you need to reach and the current observation image. Please analyze and make decisions according to the following steps:
  432. 1. Direction judgment:
  433. Based on the observation image, is the target object already in your field of vision? If it appears, in which direction is it located? For example: straight ahead/left front/right front. If it does not appear, determine in which direction it is most likely to appear.
  434. 2. Strategy decision:
  435. Based on the direction you analyzed, if there is still a certain distance from the target, what should be your next best movement strategy? You have the following options: ['move_forward', 'turn_left', 'turn_right', 'look_up', 'look_down', 'finish_task']. Please note that you only need to consider one step of strategy.
  436. If you think you are close enough (the distance is less than 1m), you should choose the 'finish_task' strategy.
  437. 3. Output format:
  438. The final answer needs to be output in a strict format. For example: **Strategy**: 'xxx' (xxx is the strategy you choose).
  439. Please analyze according to the above requirements and respond strictly in the specified format.
  440. """
  441. },
  442. {
  443. "type": "text", "text": f"target description: {text_prompt}"
  444. },
  445. {
  446. "type": "text", "text": f"observation view:"
  447. },
  448. {
  449. "type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{base64_image}"}
  450. },
  451. {
  452. "type": "text", "text": "Now please start thinking step by step"
  453. },
  454. ],
  455. }
  456. ]
  457. # text = processor.apply_chat_template(
  458. # messages, tokenize=False, add_generation_prompt=True
  459. # )
  460. # image_inputs, video_inputs = process_vision_info(messages)
  461. # inputs = processor(
  462. # text=[text],
  463. # images=image_inputs,
  464. # videos=video_inputs,
  465. # padding=True,
  466. # return_tensors="pt",
  467. # )
  468. # inputs = inputs.to("cuda")
  469. # # Inference: Generation of the output
  470. # generated_ids = model.generate(**inputs, max_new_tokens=128)
  471. # generated_ids_trimmed = [
  472. # out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
  473. # ]
  474. # output_text = processor.batch_decode(
  475. # generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
  476. # )
  477. # return output_text[0]
  478. completion = client.chat.completions.create(
  479. model="gpt-4o",
  480. timeout=500,
  481. messages=messages
  482. )
  483. return completion.choices[0].message.content
  484. def vln_subgoal_planner_with_obs(client, text_prompt):
  485. messages = [
  486. {"role": "system", "content": "You are a helpful assistant."},
  487. {"role": "user", "content": [
  488. {
  489. "type": "text", "text":
  490. """
  491. You will get a text instruction for long-distance navigation in an indoor environment.
  492. Now, your task is to decompose the text instruction into reasonable and clear sub-task goals to help the agent complete the complex navigation task step by step. All sub-tasks need to be expressed in the form of "{move to ...}", where the target in {...} can be a word or description of an object, or a description of a room area.
  493. Next, I will give you some examples:
  494. **Text prompt:**
  495. "Walk into the hallway and at the top of the stairs turn left and walk into the bedroom. Walk past the right side of the bed and turn right into the closet and all the way through to the bathroom. Stop in front of the toilet in the bathroom."
  496. **Response:**
  497. 1. Move to the {stairs at the end of the hallway}
  498. 2. Move to the {bed in the bedroom}
  499. 3. Move to the {closet}
  500. 4. Move to the {toilet in the bathroom}
  501. **Text prompt:**
  502. "Go to the wooden stairs. Go up the stairs and go between the couch and the table. Walk into the house through the sliding glass door. Go to the television. Go to the refrigerator. Go to the front of the toaster and stop"
  503. **Response:**
  504. 1. Move to the {wooden stairs}
  505. 2. Move to the {area between a couch and a table}
  506. 3. Move to the {sliding glass door}
  507. 4. Move to the {television}
  508. 5. Move to the {refrigerator}
  509. 6. Move to the {toaster}
  510. Now, please planning the following text prompt into sub-goals and respond strictly in the specified format. Do not include any other information.
  511. """
  512. },
  513. {
  514. "type": "text", "text": f"**Text prompt:** {text_prompt}"
  515. },
  516. ],
  517. }
  518. ]
  519. completion = client.chat.completions.create(
  520. model="gpt-4o",
  521. timeout=500,
  522. messages=messages
  523. )
  524. return completion.choices[0].message.content
  525. def vln_subgoal_planner_no_object(client, text_prompt):
  526. messages = [
  527. {"role": "system", "content": "You are a helpful assistant."},
  528. {"role": "user", "content": [
  529. {
  530. "type": "text", "text":
  531. """
  532. You will get a text instruction for long-distance navigation in an indoor environment.
  533. Now, your task is to decompose the text instruction into reasonable and clear sub-task goals to help the agent complete the complex navigation task step by step.
  534. Next, I will give you some examples, the sub-task must in the format of "{...}":
  535. **Text prompt:**
  536. "Walk into the hallway and at the top of the stairs turn left and walk into the bedroom. Walk past the right side of the bed and turn right into the closet and all the way through to the bathroom. Stop in front of the toilet in the bathroom."
  537. **Response:**
  538. 1. {walk into the hallway and at the top of the stairs.}
  539. 2. {turn left and walk into the bedroom.}
  540. 3. {walk past the right side of the bed}
  541. 4. {turn right into the closet and all the way through to the bathroom.}
  542. 5. {Stop in front of the toilet in the bathroom}
  543. **Text prompt:**
  544. "Go to the wooden stairs. Go up the stairs and go between the couch and the table. Walk into the house through the sliding glass door. Go to the television. Go to the refrigerator. Go to the front of the toaster and stop"
  545. **Response:**
  546. 1. {go to the wooden stairs}
  547. 2. {go up the stairs}
  548. 3. {go between the couch and the table}
  549. 4. {walk into the house through the sliding glass door}
  550. 5. {go to the television}
  551. 6. {go to the refrigerator}
  552. 7. {go to the front of the toaster and stop}
  553. Now, please planning the following text prompt into sub-goals and respond strictly in the specified format. Do not include any other information.
  554. """
  555. },
  556. {
  557. "type": "text", "text": f"**Text prompt:** {text_prompt}"
  558. },
  559. ],
  560. }
  561. ]
  562. completion = client.chat.completions.create(
  563. model="gpt-4o",
  564. timeout=500,
  565. messages=messages
  566. )
  567. return completion.choices[0].message.content
  568. def vln_anchor_planner(client, text_prompt, obss):
  569. base64_image = image_to_base64(obss)
  570. messages = [
  571. {"role": "system", "content": "You are a helpful assistant."},
  572. {"role": "user", "content": [
  573. {
  574. "type": "text", "text":
  575. """
  576. You will act as an intelligent assistant to complete a scene navigation task. You need to analyze the provided observation images and determine the most appropriate direction of movement based on the given navigation instructions. Once the direction is determined, you need to use the corresponding image to predict and describe any important objects or environmental features that you will encounter at the destination. Your description should use natural language and provide detailed descriptions of the object's external features (texture, shape, etc.). If no obvious object is found, you should describe the surrounding environment and any significant structures or markers within it.
  577. Here are the task breakdown steps you need to follow:
  578. Direction Determination: Based on the navigation instruction, you need to choose the correct directional image from the provided images. This direction should be derived from the images provided and must be a logical response to the instruction.
  579. Object Analysis: After determining the correct directional image, analyze the image to identify what "anchor objects" you will encounter at the destination. Anchor objects represent the physical objects you will encounter when following the abstract instruction. An anchor object can be any object, such as a piece of furniture, a notable landmark, or any other clearly identifiable structure. You should provide a detailed natural language description of the anchor object, including its appearance, such as shape, color, texture, and any other notable features. If no obvious object is found in the selected direction, try to describe the appearance of the destination. This might include the terrain, nearby notable structures, or any significant features.
  580. Output Response: Finally, you need to return your analysis results and provide the anchor object description.
  581. Here is an example. You must strictly follow the format of the following response:
  582. Instruction: "Walk straight through the bathroom and exit."
  583. Your Response should be in the following format:
  584. Analysis: (your analysis here)
  585. Anchor Object: e.g., A large marble sink with gold-trimmed faucets. The sink is rectangular, with polished white marble and faint gray veining running across its surface. Above the sink, there's a framed mirror with an ornate golden frame.
  586. Now, please start your analysis based on the following input information:
  587. """
  588. },
  589. {
  590. "type": "text", "text": f"**Instruction:** {text_prompt}"
  591. },
  592. {
  593. "type": "text", "text": f"**Observation images:**"
  594. }
  595. ],
  596. }
  597. ]
  598. for idx, base64_image in enumerate(base64_image):
  599. messages[1]["content"].extend([
  600. {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{base64_image}"}}
  601. ])
  602. completion = client.chat.completions.create(
  603. model="gpt-4o",
  604. timeout=500,
  605. messages=messages
  606. )
  607. return completion.choices[0].message.content
  608. def vln_anchor_planner_v2(client, text_prompt, obss):
  609. base64_image = image_to_base64(obss)
  610. messages = [
  611. {"role": "system", "content": "You are a helpful assistant."},
  612. {"role": "user", "content": [
  613. {
  614. "type": "text", "text":
  615. """
  616. You will act as an agent to complete the task of assisting scene navigation. I will provide you with a navigation instruction and a series of current observations. The navigation instruction requires the agent to move from the current position to an "object" not far away. However, the current description of the object is rough and is only a simple category description.
  617. Now, your task is to analyze the current observation and then describe the characteristics of the target object in as much detail as possible, including fine-grained elements such as appearance, shape, and texture. You may encounter two situations:
  618. 1. The target object is within the field of view of the observed image, which indicates that you can make a fine-grained description based on the observation content.
  619. 2. The target object does not exist in the field of view of the observed image. At this time, you need to associate and infer the possible characteristics of the target object as much as possible based on the indoor environment and surrounding information in the observed image, and give a description.
  620. Here is an example. You need to output the description directly without any other analysis or additional responses:
  621. Input Instruction:
  622. "Move to the marble sink in the bathroom."
  623. Your Response should like:
  624. "A large marble sink with gold-trimmed faucets. The sink is rectangular, with polished white marble and faint gray veining running across its surface. Above the sink, there's a framed mirror with an ornate golden frame."
  625. Now, please start your analysis based on the following input information:
  626. """
  627. },
  628. {
  629. "type": "text", "text": f"**Instruction:** {text_prompt}"
  630. },
  631. {
  632. "type": "text", "text": f"**Observation images:**"
  633. }
  634. ],
  635. }
  636. ]
  637. for idx, base64_image in enumerate(base64_image):
  638. messages[1]["content"].extend([
  639. {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{base64_image}"}}
  640. ])
  641. completion = client.chat.completions.create(
  642. model="o3",
  643. timeout=500,
  644. messages=messages
  645. )
  646. return completion.choices[0].message.content
  647. def EQA_generate_anchor_object(client, text_prompt):
  648. messages = [
  649. {"role": "system", "content": "You are a helpful assistant."},
  650. {"role": "user", "content": [
  651. {
  652. "type": "text", "text":
  653. """
  654. You will act as an agent to complete the task of assisting scene embodied question answering. I will provide you with a question about the scene. In order to answer this question, we must first navigate to the vicinity of the instance involved in the question.
  655. Now, your task is to analyze and determine the description of the target instance you need to move to based on the current question, which can include some necessary spatial context, such as what type of room this target instance is in and what clear objects exist around it. If you think it is difficult to infer the exact target instance for the type of question provided, please output "We need to go around and check"
  656. I will provide you with some examples. You need to output the description directly without including other analysis and additional response content.
  657. Example 1:
  658. Question: What is the white object on the wall above the TV?
  659. Response: Now, we need to go to {A TV mounted on the wall.}
  660. Example 2:
  661. Question: What is in between the two picture frames on the blue wall in the living room?
  662. Response: Now, we need to go to {A blue wall in the living room with two picture frames on it.}
  663. Example 3:
  664. Question: What should I do to cool down?
  665. Response: We need to go around and check.
  666. Your Turn:
  667. """
  668. },
  669. {
  670. "type": "text", "text": f"Question: {text_prompt} Response:"
  671. }
  672. ],
  673. }
  674. ]
  675. completion = client.chat.completions.create(
  676. model="o3-mini",
  677. timeout=500,
  678. messages=messages
  679. )
  680. return completion.choices[0].message.content
  681. def EQA_Answer_o3(client, text_prompt, obss):
  682. base64_image = image_to_base64(obss)
  683. messages = [
  684. {"role": "system", "content": "You are a helpful assistant."},
  685. {"role": "user", "content": [
  686. {
  687. "type": "text", "text":
  688. f"""
  689. You are an intelligent question-answering robot. I will ask you some questions about indoor spaces, and you must give answers.
  690. You will see a set of images collected from the same location. The location of these images is related to the question, and you must carefully reason from them to get the correct answer. If you think there is not enough information to get the answer at that location, please try to guess a close and most likely answer.
  691. Based on the user query, you must output "text" to answer the question asked by the user. Here are some examples:
  692. Example1:
  693. Question: What is to the left of the mirror?
  694. Answer: A plant in a tall vase.
  695. Example2:
  696. Question: Where is the mirror?
  697. Answer: Next to the staircase above the dark brown cabinet.
  698. Your Turn:
  699. Question: {text_prompt}
  700. """
  701. },
  702. {
  703. "type": "text", "text": f"**Observation images:**"
  704. }
  705. ],
  706. }
  707. ]
  708. for idx, base64_image in enumerate(base64_image):
  709. messages[1]["content"].extend([
  710. {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{base64_image}"}}
  711. ])
  712. completion = client.chat.completions.create(
  713. model="o3-mini",
  714. timeout=500,
  715. messages=messages
  716. )
  717. return completion.choices[0].message.content
  718. def EQA_Answer_4o(client, text_prompt, obss):
  719. base64_image = image_to_base64(obss)
  720. messages = [
  721. {"role": "system", "content": "You are a helpful assistant."},
  722. {"role": "user", "content": [
  723. {
  724. "type": "text", "text":
  725. f"""
  726. You are an intelligent question-answering robot. I will ask you some questions about indoor spaces, and you must give answers.
  727. You will see a set of images collected from the same space. These observation images is related to the question, and you must carefully reason from them to get the correct answer. Please do not respond with "I can't answer that because I didn't see..." If you think the images are not accurate enough to get an answer, please try to guess the most likely answer.
  728. Based on the user query, you must output "text" to answer the question asked by the user. Here are some examples:
  729. Example1:
  730. Question: What is to the left of the mirror?
  731. Answer: A plant in a tall vase.
  732. Example2:
  733. Question: Where is the mirror?
  734. Answer: Next to the staircase above the dark brown cabinet.
  735. Your Turn:
  736. Question: {text_prompt}
  737. """
  738. },
  739. {
  740. "type": "text", "text": f"**Observation images:**"
  741. }
  742. ],
  743. }
  744. ]
  745. for idx, base64_image in enumerate(base64_image):
  746. messages[1]["content"].extend([
  747. {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{base64_image}"}}
  748. ])
  749. completion = client.chat.completions.create(
  750. model="gpt-4o",
  751. timeout=500,
  752. messages=messages
  753. )
  754. return completion.choices[0].message.content

LLMAgent.py at commit ad4c544, no license · at the source

Overview

Authors: Shouwei Ruan1,2, Liyuan Wang1,3, Caixin Kang2, Qihui Zhu2, Songming Liu1, Xingxing Wei2, Hang Su1
  1. Department of Computer Science and Technology, Institute for AI, BNRist Center, Tsinghua-Bosch Joint ML Center, THBI Lab, Tsinghua University, Beijing, China
  2. Institute of Artificial Intelligence, Beihang University, Beijing, China
  3. Department of Psychological and Cognitive Sciences, Tsinghua University, Beijing, China
Institutions: Beihang University (China); Tsinghua University (China)
Journal: Nature communications, volume 17, issue 1, article 8061
Dates: received 13 October 2025; accepted 2 June 2026; published online 27 June 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s41467-026-74358-5 · PMID 42364983 · PMCID PMC13454140 · OpenAlex W7166300381
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), cognitive (subfield)
Methods: Graphs, Machine learning
Keywords: Computer science, Cognitive neuroscience
MeSH: Artificial Intelligence*, Brain*, Spatial Memory*, Spatial Navigation*, Cognition, Cues, Humans, Large Language Models (* major topic)
Topic: EEG and Brain-Computer Interfaces (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Citations: not cited yet (Europe PMC); 87 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repositories

Its files are read in the Code ↔ Paper reader above, with 16 matches between paragraphs and lines of code.

facebookresearch/habitat-sim

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 57ee4941dc4765240f0f91f70b2c97a919bf9038, 7 May 2026
Languages: C/C++ (196), C++ (191), Python (124), Shell (11), Jupyter (8), C (1), CUDA (1)
Size: 930 files, 532 scripts
Software Heritage: archived
Found in: “Data availability”
Holds: README, license file, environment (pyproject.toml, requirements.txt, setup.cfg, setup.py, conda-build/Dockerfile, conda-build/habitat-sim-mutex/conda_build_config.yaml), tests, continuous integration, documentation, 8 notebooks
Not found: CITATION.cff
Tools: NumPy (62 files), Pillow (11 files), Matplotlib (8 files), PyTorch (8 files), imageio (3 files), Numba (2 files), OpenCV (1 file), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
534 files

facebookresearch/habitat-lab

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 0fb6f43ffe806a8088a171b036336c093bcf604e, 7 May 2026
Languages: Python (418), Jupyter (7), Shell (6)
Size: 721 files, 431 scripts
Software Heritage: archived
Found in: “Data availability”
Holds: README, license file, environment (Dockerfile, pyproject.toml, setup.cfg, habitat-baselines/setup.py, habitat-hitl/requirements.txt, habitat-hitl/setup.py, habitat-lab/requirements.txt, habitat-lab/setup.py, scripts/habitat_dataset_processing/requirements.txt, scripts/habitat_dataset_processing/setup.py, habitat-baselines/habitat_baselines/il/requirements.txt, habitat-baselines/habitat_baselines/rl/requirements.txt), tests, continuous integration, documentation, 7 notebooks
Not found: CITATION.cff
Tools: NumPy (158 files), PyTorch (76 files), OpenCV (10 files), Matplotlib (8 files), Pillow (8 files), imageio (6 files), Numba (2 files), SciPy (2 files)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
433 files

naokiyokoyama/ovon

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 8300fcc9fcd820637ac202cb43db080229f53410, 15 May 2025
Languages: Python (82), Shell (18)
Size: 126 files, 100 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, environment (pyproject.toml, setup.py), documentation
Not found: license file, CITATION.cff, tests, continuous integration
Tools: PyTorch (41 files), NumPy (32 files), Matplotlib (6 files), seaborn (6 files), pandas (5 files), OpenCV (3 files), scikit-learn (3 files), Pillow (2 files), Hugging Face Transformers (2 files), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
101 files

XinyuSun/PSL-InstanceNav

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 7b9faa6c382c50d2920ebd129b3376ca8fbd9ea4, 6 June 2025
Languages: Python (18), Shell (2)
Size: 29 files, 20 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, environment (requirements.txt, setup.cfg, setup.py)
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (10 files), PyTorch (9 files), OpenCV (2 files), Pillow (2 files)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
21 files

facebookresearch/open-eqa

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: cfa3fce4595c1622bb2f8a38ae2ca9aae9eb685b, 20 September 2024
Languages: Python (25), Jupyter (1)
Size: 49 files, 26 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, license file, environment (requirements.txt, setup.py), 1 notebook
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (8 files), OpenCV (5 files), Pillow (4 files), imageio (2 files), Matplotlib (1 file), PyTorch (1 file), Hugging Face Transformers (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
28 files

Heathcliff-saku/BSC-Nav

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: ad4c544f7d1feb980ffad33d4eefa8ad2e39fad0, 6 November 2025
Languages: Python (466), Shell (15), Jupyter (7)
Size: 862 files, 488 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, environment (environment.yml, requirements.txt, third-party/habitat-lab/Dockerfile, third-party/habitat-lab/pyproject.toml, third-party/habitat-lab/setup.cfg, third-party/habitat-lab/habitat-baselines/setup.py, third-party/habitat-lab/habitat-hitl/setup.py, third-party/habitat-lab/habitat-lab/setup.py), tests, documentation, 7 notebooks
Not found: license file, CITATION.cff, continuous integration
Tools: NumPy (195 files), PyTorch (103 files), Pillow (26 files), OpenCV (24 files), imageio (18 files), SciPy (17 files), Matplotlib (15 files), scikit-learn (15 files), Hugging Face Transformers (10 files), h5py (2 files), NetworkX (2 files), Numba (2 files), pandas (1 file), TensorFlow (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
489 files

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41467-026-74358-5.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 6 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 1,597 scripts, each with its path and the digest of its content;
  • 16 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Code and data availability statement

The paper has a code and data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41467-026-74358-5.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 7 authors, 2 keywords, 8 MeSH terms, 27 references.

Cite

This paper

Ruan, S., Wang, L., Kang, C., Zhu, Q., Liu, S., Wei, X., & Su, H. (2026). Brain-inspired spatial intelligence for embodied agents. Nature communications, 17(1), 8061. https://doi.org/10.1038/s41467-026-74358-5

BibTeX

@article{ruan2026brain,
author = {Ruan, Shouwei and Wang, Liyuan and Kang, Caixin and Zhu, Qihui and Liu, Songming and Wei, Xingxing and Su, Hang},
title = {{Brain-inspired spatial intelligence for embodied agents}},
journal = {Nature communications},
year = {2026},
month = jun,
volume = {17},
number = {1},
pages = {8061},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/s41467-026-74358-5},
url = {https://doi.org/10.1038/s41467-026-74358-5},
pmid = {42364983},
pmcid = {PMC13454140}
}

RIS

TY - JOUR
AU - Ruan, Shouwei
AU - Wang, Liyuan
AU - Kang, Caixin
AU - Zhu, Qihui
AU - Liu, Songming
AU - Wei, Xingxing
AU - Su, Hang
TI - Brain-inspired spatial intelligence for embodied agents
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/06/27
VL - 17
IS - 1
SP - 8061
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/s41467-026-74358-5
UR - https://doi.org/10.1038/s41467-026-74358-5
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41467-026-74358-5",
"type": "article-journal",
"title": "Brain-inspired spatial intelligence for embodied agents",
"container-title": "Nature communications",
"author": [
{
"family": "Ruan",
"given": "Shouwei"
},
{
"family": "Wang",
"given": "Liyuan"
},
{
"family": "Kang",
"given": "Caixin"
},
{
"family": "Zhu",
"given": "Qihui"
},
{
"family": "Liu",
"given": "Songming"
},
{
"family": "Wei",
"given": "Xingxing"
},
{
"family": "Su",
"given": "Hang"
}
],
"container-title-short": "Nat Commun",
"volume": "17",
"issue": "1",
"page": "8061",
"DOI": "10.1038/s41467-026-74358-5",
"PMID": "42364983",
"PMCID": "PMC13454140",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41467-026-74358-5",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
27
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41598-026-43529-1 [code]
A spiking neural network inspired by neuroscience and psychology for Western mode- and key-conditioned music learning and composition.
Journal: Scientific reports
In common: imageio, TensorFlow, NetworkX, 9 other tools, cognitive, 1 reference
[2] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: imageio, Numba, TensorFlow, 10 other tools
[3] doi:10.1371/journal.pcbi.1014571 [code]
SynAPSeg: A novel dataset and image analysis framework for deep learning-based synapse detection and quantification.
Journal: PLoS computational biology
In common: imageio, Numba, TensorFlow, 10 other tools
[4] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: Hugging Face Transformers, Numba, TensorFlow, 10 other tools
[5] doi:10.1038/s41598-026-57519-w [code]
Automated segmentation of neurons and spinal cord structures in immunofluorescence images using SpineDL.
Journal: Scientific reports
In common: imageio, Numba, NetworkX, 9 other tools
[6] doi: [code]
Real-time closed-loop feedback system for mouse mesoscale cortical signal and movement control
Journal: eLife
In common: imageio, Numba, TensorFlow, 9 other tools
[7] doi:10.1364/boe.605322 [code]
Generalized plaque digitization framework for multi-dimensional mesoscopic images.
Journal: Biomedical optics express
In common: imageio, TensorFlow, OpenCV, 9 other tools
[8] doi:10.7554/elife.110588 [code]
Opening the black box toward a modular approach to spike sorting.
Journal: eLife
In common: Numba, TensorFlow, NetworkX, 9 other tools
[9] doi:10.1038/s41467-026-74357-6 [code]
Hippocampo-neocortical interaction as compressive retrieval-augmented generation.
Journal: Nature communications
In common: Hugging Face Transformers, NetworkX, Pillow, 7 other tools, 2 references
[10] doi:10.3389/fnsys.2026.1822122 [code]
Convergence-divergence circuits for multimodal integration of innate and learned opponent valences.
Journal: Frontiers in systems neuroscience
In common: Numba, TensorFlow, NetworkX, 9 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.