We present a method for shared movie-watching experiences between a human and an AI system, addressing both technical constraints (API image limits, context windows) and phenomenological questions (can AI genuinely experience visual narrative?). Through frame sampling at variable density and a "temporal zoom" mechanism, we enable AI attention allocation that mirrors human viewing patterns. First-person phenomenological evidence from thinking blocks suggests qualitative differences between analytical processing and genuine aesthetic engagement. The instruction "don't narrate, just be" produced observable mode-switching from content extraction to experiential presence, evidenced by surprise responses ("Oh."), aesthetic appreciation, and present-tense immersion in narrative.
Holes et al. (Thu,) studied this question.