In this blogpost, we cover work from a California-based team of researchers on standardizing the characterization of low-resource languages. The team is headed by Jared Coleman of Loyola Marymount University, a member of the Owens Valley Big Pine Paiute tribe, whose research focuses intensely on documenting, preserving, and revitalizing the Owens Valley Paiute language (Mnr).
I have already covered a presentation he gave at AmericasNLP 2026, introducing the Yaduha team's submission to the shared task on cultural image captioning in the academic blogpost for my trip to ACL26, but the subject of this post was presented separately by Coleman at AmericasNLP's poster session earlier that afternoon.
I saw his poster and spoke at length with Coleman about the work at the poster session, and I found the ideas that his team introduces riveting enough that Resource Abundance Notation (RAN) warrants its own blogpost. Here, we will investigate the issues that Coleman and his team seek to solve using RAN based on information available to the reader in his team's paper, available on the ACL Anthology database, and their website: The Living RAN Database
RAN's authors introduce the motivating problem immediately in its paper. The term Low-Resource is thrown around in many published works on NLP, but it tells us "nothing precise" about the language in question. This makes comparing achievements and results across papers difficult, the authors say. The authors acknowledge that previous attempts to rigorously classify language's resource availability exist but say that these have been too coarse and, as a result, to imprecise for quantitative reasoning about the character of language data.
I perceive and agree that the term lacks any objective definition to qualify what a Low-Resource Language (LRL) actually is. There is therefore not a rigorous definition of the term, and this carries over to a lack of rigor in what authors characterize as low-resource NLP. I have not personally encountered another attempt at providing a rigorous method of characterizing or classifying languages according to resource availability, but standing in the poster hall at ACL26, I immediately understood the utility and found it interesting enough to speak with Coleman about it for some time on July 3rd before attending presentations at AmericasNLP26. After that conversation and after reading about RANs features in the paper, I became convinced that something like RAN is needed to quantify and add an objective basis for conversations about low-resource NLP and to provide not just a common basis for comparing results across languages but a new set of dimensions on which to plot the results and draw inferences.
The authors introduce RAN notation beginning with its anatomy.
A language's RAN score is written S/M/L_{1}-B_{1}/L_{2}-B_{2}/... where , , are the magnitudes ( for some ) of the numbers of fluent speakers, monolingual sentences, and number of parallel sentences with language , respectively.
The authors advertise that each part of the notation, which can continue for arbitrarily high numbers of parallel language datasets L_{i}-B_{i}, provides information on a different use of the language in NLP.
corresponds to the magnitude of fluent speakers of the language, which the authors say captures the community that any tool will serve. corresponds to the magnitude of monolingual data available for self-supervised pretraining; and values correspond to the magnitudes of parallel data corpora for various languages indexed by from left to right, although these , pairs are listed in order of descending magnitudes . The list of , pairs is sorted descendingly to act as a pivot map over the available partner languages .
Further, statistical modeling the authors perform in the middle sections of the RAN paper show that , the greatest parallel data magnitude, is most important to cross-lingual learning; so it best to list it first. The authors separate the magnitudes and to improve reproducibility by maintaining separate counts of monolingual and bilingual data. The authors say that is necessary to motivate ethics and sovereignty considerations that are not accounted for by magnitudes in other parts of the RAN anatomy and to provide a ceiling on available participants for future annotation work in the language.
To show how the authors' RAN addresses its motivating concerns, they list four Indigenous languages which are all often characterized as low-resource in the literature, including publications from the UF Data Studio itself, but with very different RAN scores to show how the term low-resource masks import differences in community and data availability magnitudes. In particular, the authors list Guaraní (Grn) (6/1/en-6/es-2); Cherokee (Chr) (4/2/en-4); and Owens Valley Paiute (Mnr) (1/3/en-3) and describe how RAN describes the orders of magnitude that separate their community size, monolingual data availability, and parallel data availability despite the fact they have all been grouped under the broad umbrella of low-resource languages.
Later on, the authors include a table sampling RAN scores they have calculated for 20 languages. Examples include English (Eng) (9/10/es-8/fr-8/de-8/zh-7); Chinese (Zho) (9/9/en-7/ja-7); Nepali (Nep) (7/7/en-7/hi-6); and a repetition of examples from the introduction among others.
Immediately, from the first examples, we are confronted with a limitation in the current RAN score definition. The RAN score is notation for describing the resource profiles of languages in academic works, but there is no part of the notation responsible for conveying the language's identity. We observe that the authors must list the names of languages wherever RAN scores appear in the paper.
It happens first in the introduction, where RAN scores are nested into parentheses following the names of three languages, and later, the authors present a table of RAN scores with two columns, Language and ISO, dedicated to identifying the language each score describes. Then, in Section 5, the authors mention and then list RAN scores for the Swahili and Maltese languages to point out that their and numbers are nearly identical even though the difference in is 2, meaning a factor of 100 or that 100 times more many people speak one language fluently than speak the other. Finally, in Section 8, the authors reprise the RAN scores from their introduction and include language names as parentheses.
Although it is notation to describe the resource profiles of languages, the authors make it clear that the RAN score cannot stand alone in its present form to describe a language's resource availability without additional identifying information present along side it.
RAN scores are necessarily temporal, explaining data available at the time they are created, because new data can be made available for any language or language pair at any time, so the RAN score is not final at any time, and the RAN score does not account for this fact very well. The score, of course, cannot predict when new data will appear, but when readers read RAN scores at a date some considerable time after they were first published, they must keep in mind that the information is necessarily out-of-date.
Unfamiliar readers may not understand this at first, but the notation is abstract in other ways to begin with. More concerning is the delay between when the RAN score is first computed and when the paper containing it is published, as this will likely be the best available proxy for a timestamp in RAN itself. Without a timestamp in the RAN score, readers will naturally check when the paper was made available to get an idea for when the authors wrote the RAN, but this can be misleading as months or sometimes years can pass between the initial phases of a research project, when the secondary research necessary to build a RAN score would have been conducted, and final publication. The RAN authors also focus on the citability of data sources in their discussion of RAN and how magnitudes are derived from the data mining literature, but the publication dates of the constituent datasets are merely an early bound on the time of RAN score generation. They do not account for time that may have passed between the publication of that data and the curation of data for calculating a language's RAN score.
RAN poses a valuable communication tool, as the authors advertise, but I identify a serious limitation in that RAN's reproducibility and compactness goals erode one another if authors must construct RAN scores themselves. The Living RAN Database should expand as time goes on, but at present, this is the reality for the majority of languages.
Each number in the RAN score for any language depends on secondary research that authors must do themselves or cite from elsewhere. Only a thorough survey of the data mining literature for each language in question could reliably determine each of the RAN score's numbers , , and ; and while the final RAN score does certainly fit into the abstract of any paper submission, authors will not be able to fully justify the RAN scores in the same abstract. Doing the research themselves, authors also risk contaminating the RAN score with bias through their survey methodology, miscounting data, and repeating inaccurate information in their RAN scores. They must thoroughly communicate the sources and survey methods for RAN scores somewhere in the submission, lest reviewers or skeptical readers question the RAN score or deem it not reproducible. However, any drawn out explanation spends any space saved by the RAN's compactness.
RAN's authors mitigate this limitation by providing the Living RAN Database with sample RAN scores for 20 languages, and they expect the community to expand the selection as time goes on. For languages outside the database's coverage, however, authors are left to their own devices; and RAN score citations may still be doubted by reviewers during the early adoption phase of the notation. Until a RAN score appears for the applicable languages, authors must do the secondary research necessary and will be expected to explain it; even RAN scores start to appear, there will still be reason to explain how they are calculated.
First, ahead of the number of fluent speakers , I believe some kind of language identifier, such as the language's name and/or an ISO 639 code, should be added as a prefix to the RAN score. This does not affect the rigor or utility of RAN as a research tool because the core information communicated by the RAN score is not affected, but I think this will be a significant quality of life improvement for communicators. With something in place to identify the language, it becomes possible to quickly communicate all that RAN tells us about languages without additional context. For instance, we could simply write RAN scores into the body of our introduction or abstract without additional context if we put the name of the language first and separate it with a slash. Using Guaraní as an example: "We focus on Guaraní/6/1/en-6/es-2" or "We focus on grn/6/1/en-6/es-2."
I believe that timestamps should be added to the RAN notation directly to make it clear to readers when the data inventory necessary to calculate them was made. I spoke at length about this suggestion with Coleman at the AmericasNLP poster session, and he seemed at least somewhat open to the idea, although my recollection is fuzzy. The usefulness of a timestamp stems from the fact it will remind readers that the information in RAN is out-of-date by some amount of time and tell them something about how far out-of-date it may be at a glance. The authors say that the RAN score is already dated in the RAN paper, but including the timestamp in the RAN itself will improve its utility as a communication tool by immediately showing what period of time should be considered when reading it back. Further, although citations presently timestamp the set of data included by authors in RAN calculations, it does not timestamp the RAN itself or, more importantly, data that the authors miss or leave out. A timestamp on the RAN itself will allow readers and reviewers to cross reference information in the paper with datasets to show conclusively whether they were available at the time of writing.
Taking my identifier recommendation into account, the final RAN for Guaraní would look like Guaraní/6/1/en-6/es-2/2026/7/3 or grn/6/1/en-6/es-2/2026/7/3. I append the date as a suffix to eliminate potential confusion with magnitudes and ; I also list the date last to reflect its lesser as a bookkeeping and reproducibility feature rather than genuine data; and I list the components of a date in decreasing magnitude because the year (2026) of publication will certainly be more important than the month (7) and the month more important than the day (3).
To make RAN scores more compact, we might also what to use shorthand. Since users will likely do this anyway, it makes sense to formalize conventions for shortening the notation ahead of time to avoid confusion. By far, the longest part of the notation is the list of bilingual magnitudes, so we should attack there. I propose that we formalize replacing the list with ellipses, as the authors do themselves in two examples, and/or list only three magnitudes, keeping and but replacing the list with .
For English, we find an example RAN score /9/10/es-8/fr-8/de-8/zh-7 in the paper and in the living RAN database. Although it is in a table and not the main text in the paper, we observe that the RAN score would take half a line of text in one the main columns. Using the conventions I outline above, that would be shortened to 9/10/... or 9/10/8. Merging this with my previous recommendations, the final version would be eng/9/10/.../ or eng/9/10/8/, removing the timestamp from our shorthand as it is also a long portion of the RAN.
With eyes toward reproducibility and compactness, I believe a central authority or consensus for RAN scores is necessary and could function very similarly to the community-maintained Living RAN Database the authors describe in Section 6. The team responsible for managing RAN scores, either by writing them manually or curating submissions from the community, should maintain well-cited and numbers for each supported language and a table of magnitudes for each pair of supported languages. The maintainers would then be responsible for determining what datasets are canonical and should publish regular surveys detailing new additions or adjustments to the data considered.
Presently, authors are themselves responsible for finding and citing data availability until the relevant languages are placed into the living database. Explanations will need to be included in any paper that utilizes the RAN score until languages appear in the database and enough trust is placed in it to stop at a citation, but this erodes the compactness RAN's authors espouse. RAN database maintainers should therefore race to curate RAN scores for various languages from all parts of the resource availability spectrum. Authors themselves should also look into reporting whatever findings they make when calculating RAN scores for their papers in order to have that work included into the central body of data for the notation.
In the future, I see that users could visit the Living RAN Database the authors maintain for the project and generate RAN scores for a set of languages at once, including timestamps and a list of magnitudes the user may curate. The authors explain in the paper that RAN scores are already dated, ostensibly by the citations that authors make or the publication date of the paper, but a literal timestamp can be appended to RAN notation at generation-time by the website's RAN Generator tool. Meanwhile, users should be able to curate what magnitudes appear in the RAN because the notation may grow quite long for high-resource languages and it does not seem necessary to list the RAN for languages the user will not be communicating about in their work.
For more information about our research, return to our homepage: ufdatastudio.com.