Skip to main content

Accessible AI Research

Making picture-book stories fully accessible.

BlindStoryFlow is a narrative-aware AI system that preserves characters, plot flow, and emotional context across picture-book pages for blind and low-vision readers.

Narrative output

“The same child from page 1 continues her search in the snowy forest, still carrying the blue lantern. Her expression is more determined than afraid.”

Continuity memory active

300+

Annotated pages

8.4%

Unsupported claim rate

86%

Comprehension score

5.1%

Plot continuity error rate

The problem

Screen readers read words. Stories also live in pictures.

Illustrations carry facial expressions, character actions, objects, and emotional cues that standard screen readers do not interpret. Existing image-description tools often treat every page as an isolated scene.

Missing context

Important visual information is absent from the reading experience.

Broken continuity

Characters and objects can be described inconsistently from page to page.

Object lists

Descriptions identify items but often fail to explain their role in the story.

Example

From isolated description to narrative context

Standard AI

“A child and a dog in snow. Trees in the background.”

BlindStoryFlow

“The same child from page 1 continues searching the snowy forest, still carrying the blue lantern she found earlier. Her dog, missing since page 3, has not yet reappeared.”

Our Research

Multimodal AI for accessible picture-book narration

Our research treats picture-book accessibility as a narrative-integration problem: preserve the author's original text while adding concise, grounded visual information that helps blind and low-vision readers understand illustrations, character actions, emotion cues, and story continuity.

Peer-reviewed recognition: Our research was accepted to and presented at the International Symposium of Computer Science and Educational Technology (ISCSET) 2026.

What the research contributes

auto_storiesNarrative integration

The system goes beyond standalone captions by combining printed text and illustration evidence while preserving wording, pacing, dialogue, and cross-page meaning.

device_hubContinuity-aware pipeline

A shared story-context memory tracks character identities, recent events, and critical visual details so recurring characters and plot states remain consistent.

verifiedReliability safeguards

Word budgets, dialogue protection, original-text restoration, repetition control, unsupported-language filtering, and safety fallbacks constrain the generated narration.

Method

Research workflow

  • Classify the page and extract printed text

    Each page image is processed to identify page type and recover readable story text, dialogue, captions, signs, and labels without summarizing or paraphrasing.

  • Generate a grounded visual description

    The vision model describes only observable details such as expressions, positioning, objects, clothing, background, colors, and spatial relationships.

  • Update story-context memory

    The system records character identities, confidence, recent events, prior context, and critical visual cues to reduce repetition and prevent cross-page confusion.

  • Generate accessible narration

    The original text forms the narrative spine, with concise factual visual details added under identity, continuity, anti-repetition, and length constraints.

  • Clean and safeguard the output

    Structural and rule-based post-processing restores altered text, protects dialogue, removes unsupported sensory or emotional language, normalizes punctuation, and falls back to the original text when needed.

Validation

Automated testing and reader preference study

Evaluation covered 61 pages from four books and a within-subjects study with 16 adult blind and low-vision participants comparing original and visually enriched narration.

Overall preference

76.6%

Of all 512 judgments, 392 selected enriched narration, 55 selected the original, and 65 were neutral.

Sided preference

87.7%

Among 447 non-neutral judgments, enriched narration was preferred, with a 95% confidence interval of 84.3%–90.4%.

Participant-level result

15 of 16

Participants cast a majority of their sided votes for enriched narration; the participant-level sign test gave p = 0.000519.

Preference by dimension

Better descriptions: 94.2%

Recommend for readers: 92.7%

More enjoyable: 86.6%

Easier to remember: 76.0%

The enriched version was preferred across all eight excerpts and all four questions.

Automated quality results

Original Text Preservation Rate: 1.00

Dialogue Preservation Rate: 1.00

Human-verified unsupported claims: 0.0%

Human-verified continuity errors: 0.0%

The mean word expansion ratio was 1.36 across the 61 evaluated pages.

Study design
Participants

Sixteen adults who self-reported visual impairment: five blind from birth, six with acquired blindness, and five with low vision.

Materials

Eight excerpts from three books, evaluated on description quality, enjoyment, memorability, and willingness to recommend.

Accessibility and oversight

The final survey was accessibility-corrected and counterbalanced. Vista Center reviewed and approved the study under its organizational research guidelines.

How to interpret the findings

The results provide evidence of feasibility and preference, not population-level effectiveness. The sample was small, all participants were adults, and the study measured preference rather than controlled comprehension. Evaluation with children, larger samples, independent human review, broader book formats, and comparisons with captioning, alt text, and human-authored descriptions remain future work.

Team & Dataset

Independent student research

The project combines multimodal AI, accessibility research, dataset development, and evaluation with blind and low-vision readers.

Research team

Aarush Jain

Lead researcher · System architecture · Evaluation

Designed and implemented the narrative-aware system, developed the evaluation methodology, and led the research study.

Gayatri Bhimaraju

Co-researcher · Dataset · Annotation

Applied accessibility guidelines to help create the gold-standard narration dataset and co-designed the experimental evaluation framework.

Vista Center for the Blind and Visually Impaired

Accessibility partner · Participant engagement

Supported engagement with blind and low-vision participants and provided accessibility-focused input for the research evaluation.

Dataset structure

TierPagesPurpose
Core300Primary benchmark and evaluation
Extended500Broader illustration and narrative diversity
Large800Future scaling and external research use
Sources

Bloom Library, African Storybook, and Free Kids Books.

Annotation

Literal content, story function, and cross-page continuity.

Use

Academic and accessibility research subject to request review.

Dataset Access

Request access to the research dataset

Provide your contact information and affiliation. The request will be emailed to the BlindStoryFlow research team.