Reflection Blogging: Module 5 (Hyperlinked Communities)

As I was reviewing our Module 5 readings, I re-watched the American Library Association (ALA)’s YouTube video about equity at the Multnomah County Library in Portland, Oregon, and was surprised by the “quality” of YouTube’s auto-generated closed captions.

For Assignment X, I retroactively added captions to my YouTube video after reading INFO287’s guidelines for media-based assignments. Always a writer/editor, I reviewed the Canva-generated captions and edited them as necessary. The video was about Scandinavian libraries, so naturally I had to correct proper nouns like “Dokk1” and “Oodi” and Latin phrases such as genius loci.

At 1:13, LeFoster Williams, a Black Cultural Library Advocate and Library Assistant at Multnomah County Library, says “patrons” two times; YouTube’s captions read as “pictures” and “pages” instead. At 3:45, Mrs. Tang, a patron, is speaking with Sean Khoo, a librarian at Multnomah County Library, about their bilingual services. Mrs. Tang, speaking Cantonese, is captioned as saying “they kill the whole dog at home.” The error was distracting and could be culturally damaging if read by an uninformed, English-speaking YouTube user. Meanwhile, the auto-generated captions accurately capture what the white Director of Libraries for the Multnomah County Library System is saying throughout.

Naturally, I did some research on the history of closed captioning and the limits of automatic speech recognition technology. The National Captioning Institute, a nonprofit, was created in 1979 to expand television access to the deaf and hard of hearing, and eventually non-native English speakers. Now, it is a model of inclusivity in media consumption (Burchett, 2019 & National Captioning Institute, 2020-a). Live, or real-time, captioning has an accuracy rate of 98% (National Captioning Institute, 2020-b). According to the International Speech Communication Association (ISCA), with manual keyboard input, human transcribers can achieve a near 100% accuracy rate at a lag rate of 8 seconds maximum (ISCA, 2024).

File:Closed captioning symbol.svg

Both commercial and open-source automatic speech recognition (ASR) is error-laden, with open-source ASR producing more errors (Russell, 2024). It is a known fact that ASR has a negative performance bias against non-native speakers or those with dialects (Russell, 2024). Outrageously, a study conducted by the Stanford Computational Policy Lab showed ASR misunderstands Black male speakers at twice the rate it misunderstands white male speakers (Koenecke, 2020).  You can explore infographic study results at this Stanford site.

It is still recommended that humans manually correct an ASR-produced transcript, as I did with my captions made in Canva (Russell, 2024). Thankfully, the ALA upheld its core value of accessibility with accurate captioning in this video.

Up next… voice AI “accent conversion” technology from Krisp (Krisp, 2026).

References

American Library Association. (2019). Multnomah County Library: Creating conditions for equity to flourish [Video]. YouTube. https://www.youtube.com/watch?v=SKGlxh-zc0Y 

American Library Association. (2024, January). Core values of librarianship. https://www.ala.org/advocacy/advocacy/intfreedom/corevalues

Burchett, M.H. (2019). Closed captioning developed. EBSCO. https://www.ebsco.com/research-starters/communication-and-mass-media/closed-captioning-developed/

Hewitt, J., Bateman, A., Lambourne, A., Ariyaeeinia, A., Sivakumaran, P. (2024). Real-time speech-generated subtitles: problems and solutions. International Speech Communication Association, https://www.isca-archive.org/icslp_2000/hewitt00_icslp.html
INFO 287 – The Hyperlinked Library. (2026). Guidelines for media-based assignments. https://287.hyperlib.sjsu.edu/assignments/guidelines-for-media-based-assignments/

 

Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., Toups, C., Rickford, J.R., Jurafsky, D., & Goel, S. (2020, March 23). Racial disparities in automated speech recognition. Proceedings of the National Academy of Sciences of the United States of America, 117(14), 7684–7689. https://doi.org/10.1073/pnas.1915768117

Krisp Technologies, Inc. (2026). AI accent conversion. https://krisp.ai/ai-accent-conversion/

National Captioning Institute. (2020-a). About us. https://www.ncicap.org/about-us

National Captioning Institute. (2020-b). Viewer FAQ. https://www.ncicap.org/viewer-faq

Leave a Reply

The act of commenting on this site is an opt-in action and San Jose State University may not be held liable for the information provided by participating in the activity.

Your email address will not be published. Required fields are marked *