University of Florida Data Studio 🐊

The UF Data Studio is home for research projects for all parts of the Data Pipeline. We are passionate about developing new ways humans can interact with their data and understanding the fundamental research issues all along the data pipeline.

Avatar image
Latest Updates
Figure 1 from Zhang et al. (ACL 2026), showing the LoopTool closed-loop pipeline for iterative tool-use enhancement.

LoopTool — When Tool-Use Training Needs a Closed Loop

Aug 1, 2026
A blogpost on an ACL 2026 paper that couples GRPO training with greedy probing, judge-guided label repair, and error-driven data expansion for robust LLM function calling.
Highrise buildings standing on the left of streetlights and streetlamps with the overhanging roof and tall glass facade of the Colorado Convention Center looming to the right.

RAN Notation - Introducing Objectivity to Descriptions of Resource Scarcity in NLP

Jul 31, 2026
We discuss a recent breakthrough in language documentation. Researchers introduce a new formalism to concisely communicate the number of speakers and corpus sizes of languages.
High rise buildings standing on the left of streetlights and streetlamps with the overhanging roof and tall glass facade of the Colorado Convention Center looming to the right.

The OpenMantra Dataset for Multimodal Manga Translation

Jul 31, 2026
OpenMantra is a dataset from the same lab behind Manga109, published in 2021. In this post, we contrast OpenMantra with the much more well known Manga109, and we discuss how the UF Data Studio may use it in the future.
Figure 1 from Feng and Yang, comparing FFT and LoRA training curves on Quora paraphrase detection.

When DPO and LoRA Stop Helping — Lessons from GPT-2 Scale

Jul 30, 2026
A blogpost on an empirical study of SFT, DPO, FFT, and LoRA on a 124M-parameter GPT-2 backbone for paraphrase detection and Shakespeare sonnet continuation.

System for Future Reference Patterns

Jul 24, 2026
Detecting expressions which refer to future events, predicting the future....
Figure 1 from Elbouanani et al. (ACL 2026 Findings), showing entity-level audit scores for countries aggregated across tasks, languages, models, and prompts.

Auditing LLM Bias at Billion-Point Scale

Jul 23, 2026
A blogpost on an ACL 2026 Findings paper that uses named entities as probes to measure structural bias in large language models across politics, geopolitics, and finance.
View all posts →

Proudly Funded By

© Copyright 2026 by UF Data Studio. Built with ♥ by ceg.me. [trailers]