Publications and Research

Document Type

Article

Publication Date

8-13-2026

Abstract

This article presents the Russian Language-Monolingual corpus (RusLan-M, v.1.0), a longitudinal multimedia collection of early child speech from two Russian-speaking monolingual children: Tosya (ages 0;10–3;10, 246 recordings) and Yasha (ages 1;04–3;00, 42 recordings). The corpus consists of approximately 41 h (2,454 min.) of video recordings and 35,386 child utterances, available with transcriptions in the CHAT format on TalkBank. The corpus adheres to strict ethics requirements for data sharing with anonymization. We also conducted two exploratory investigations of the acquisition of Russian morphology using the mean length of utterance (MLU) and the newly developed Index of Productive Syntax for evaluating grammatical complexity in the nominal system of Russian (IPSyn-NP-R). These investigations illustrate how a comprehensive analysis of syntactic and morphological structures in Russian language development can be conducted and highlight the kinds of questions that can be answered using the RusLan-M data. The RusLan-M corpus represents an application of corpus linguistics methods to the Russian language and fills a significant gap in the limited data and resources available for Russian child language research.

Share

COinS
 
 

To view the content in your browser, please download Adobe Reader or, alternately,
you may Download the file to your hard drive.

NOTE: The latest versions of Adobe Reader do not support viewing PDF files within Firefox on Mac OS and if you are using a modern (Intel) Mac, there is no official plugin for viewing PDF files within the browser window.