Shmuel Berman

Shmuel Berman

Princeton University

New York, New York

Substack

Github

LinkedIn

Goodreads

About Me

I'm a 3rd year PhD student at Princeton University advised by Jia Deng, and an Anthropic Fellow. I received my B.S in Computer Science from Columbia University in 2024, where I was fortunate enough to do research with Professors Mark Santolucito, Kathleen McKeown, and Baishakhi Ray.

I love hiking, horror movies, peanut butter, and reading. Good reasons to contact me: research, book recs, hiking recs.

Research Interests

I study perception, memory, and reasoning in foundation models, with a particular interest in the visual and embodied capabilities required for robust interaction with the physical world. More broadly, I care about how these systems perceive, remember, and act in non-textual environments, including robotic settings.

Selected Work

Figure 1 from Good Memory Has ECC, contrasting two memory systems on efficiency, compression, and calibration.

Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy

Preprint, 2026

Shmuel Berman, Jia Deng

ECCBench evaluates memory beyond raw accuracy along three axes: efficiency (the FLOPs needed to answer from memory), compression (whether compressible inputs are remembered better), and calibration (whether a system abstains when it is uncertain). Pretrained VLMs compress over text but not video, and are poorly calibrated on both.

Chart from Claude plays robotics showing per-model embodiment scores broken down by control interface.

Claude plays robotics

Anthropic, 2026

Shmuel Berman, Michael Ilie, Jia Deng, Daniel Freeman

An evaluation of whether frontier language models can perceive a scene, track a robot’s state, and issue actions that reliably change the physical world, across classic control tasks, simulated quadrupeds and humanoids, a robotic arm, and a real Unitree Go2.

Figure from VLMs have Tunnel Vision showing the nonlocal visual reasoning tasks.

VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs

NeurIPS 2025 Spotlight

Shmuel Berman, Jia Deng

An evaluation suite for nonlocal visual reasoning in leading vision-language models, covering comparative perception, saccadic search, and smooth visual search.

Figure from the zebra puzzle paper showing an example puzzle layout and clues.

Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems

Shmuel Berman, Kathleen McKeown, Baishakhi Ray

A multi-agent framework that combines language models with a theorem prover to translate natural-language clues into structured constraints and solve zebra puzzles more reliably.