Accessible AI Research
Making picture-book stories fully accessible.
BlindStoryFlow is a narrative-aware AI system that preserves characters, plot flow, and emotional context across picture-book pages for blind and low-vision readers.
“The same child from page 1 continues her search in the snowy forest, still carrying the blue lantern. Her expression is more determined than afraid.”
300+
Annotated pages
8.4%
Unsupported claim rate
86%
Comprehension score
5.1%
Plot continuity error rate
The problem
Screen readers read words. Stories also live in pictures.
Illustrations carry facial expressions, character actions, objects, and emotional cues that standard screen readers do not interpret. Existing image-description tools often treat every page as an isolated scene.
Important visual information is absent from the reading experience.
Characters and objects can be described inconsistently from page to page.
Descriptions identify items but often fail to explain their role in the story.
Example
From isolated description to narrative context
“A child and a dog in snow. Trees in the background.”
“The same child from page 1 continues searching the snowy forest, still carrying the blue lantern she found earlier. Her dog, missing since page 3, has not yet reappeared.”
Our Research
Multimodal AI for accessible picture-book narration
Our research treats picture-book accessibility as a narrative-integration problem: preserve the author's original text while adding concise, grounded visual information that helps blind and low-vision readers understand illustrations, character actions, emotion cues, and story continuity.
What the research contributes
The system goes beyond standalone captions by combining printed text and illustration evidence while preserving wording, pacing, dialogue, and cross-page meaning.
A shared story-context memory tracks character identities, recent events, and critical visual details so recurring characters and plot states remain consistent.
Word budgets, dialogue protection, original-text restoration, repetition control, unsupported-language filtering, and safety fallbacks constrain the generated narration.
Method
Research workflow
- Classify the page and extract printed text
Each page image is processed to identify page type and recover readable story text, dialogue, captions, signs, and labels without summarizing or paraphrasing.
- Generate a grounded visual description
The vision model describes only observable details such as expressions, positioning, objects, clothing, background, colors, and spatial relationships.
- Update story-context memory
The system records character identities, confidence, recent events, prior context, and critical visual cues to reduce repetition and prevent cross-page confusion.
- Generate accessible narration
The original text forms the narrative spine, with concise factual visual details added under identity, continuity, anti-repetition, and length constraints.
- Clean and safeguard the output
Structural and rule-based post-processing restores altered text, protects dialogue, removes unsupported sensory or emotional language, normalizes punctuation, and falls back to the original text when needed.
Validation
Automated testing and reader preference study
Evaluation covered 61 pages from four books and a within-subjects study with 16 adult blind and low-vision participants comparing original and visually enriched narration.
76.6%
Of all 512 judgments, 392 selected enriched narration, 55 selected the original, and 65 were neutral.
87.7%
Among 447 non-neutral judgments, enriched narration was preferred, with a 95% confidence interval of 84.3%–90.4%.
15 of 16
Participants cast a majority of their sided votes for enriched narration; the participant-level sign test gave p = 0.000519.
Better descriptions: 94.2%
Recommend for readers: 92.7%
More enjoyable: 86.6%
Easier to remember: 76.0%
The enriched version was preferred across all eight excerpts and all four questions.
Original Text Preservation Rate: 1.00
Dialogue Preservation Rate: 1.00
Human-verified unsupported claims: 0.0%
Human-verified continuity errors: 0.0%
The mean word expansion ratio was 1.36 across the 61 evaluated pages.
Sixteen adults who self-reported visual impairment: five blind from birth, six with acquired blindness, and five with low vision.
Eight excerpts from three books, evaluated on description quality, enjoyment, memorability, and willingness to recommend.
The final survey was accessibility-corrected and counterbalanced. Vista Center reviewed and approved the study under its organizational research guidelines.
The results provide evidence of feasibility and preference, not population-level effectiveness. The sample was small, all participants were adults, and the study measured preference rather than controlled comprehension. Evaluation with children, larger samples, independent human review, broader book formats, and comparisons with captioning, alt text, and human-authored descriptions remain future work.
Team & Dataset
Independent student research
The project combines multimodal AI, accessibility research, dataset development, and evaluation with blind and low-vision readers.
Research team
Lead researcher · System architecture · Evaluation
Designed and implemented the narrative-aware system, developed the evaluation methodology, and led the research study.
Co-researcher · Dataset · Annotation
Applied accessibility guidelines to help create the gold-standard narration dataset and co-designed the experimental evaluation framework.
Accessibility partner · Participant engagement
Supported engagement with blind and low-vision participants and provided accessibility-focused input for the research evaluation.
Dataset structure
| Tier | Pages | Purpose |
|---|---|---|
| Core | 300 | Primary benchmark and evaluation |
| Extended | 500 | Broader illustration and narrative diversity |
| Large | 800 | Future scaling and external research use |
Bloom Library, African Storybook, and Free Kids Books.
Literal content, story function, and cross-page continuity.
Academic and accessibility research subject to request review.
Dataset Access
Request access to the research dataset
Provide your contact information and affiliation. The request will be emailed to the BlindStoryFlow research team.