Research Interests: Language Models, Pre-training, Generalization, LLM Evaluation, Mathematical Reasoning, Interpretability, Parameter-Efficient Fine-Tuning
I’m a Computer Science PhD candidate at the University of Massachusetts Lowell, advised by Prof. Anna Rumshisky. My research takes a data-centric approach to understanding language-model abilities: how training data can shape which capabilities emerge, how data exposure can reveal when models move from memorization toward generalization, and how evaluation data can uncover failures in LLMs that standard benchmarks miss. My work spans language-model pre-training, training dynamics, generalization, model analysis and interpretability, and LLM evaluation. I’m interested in Research Scientist and Applied Scientist roles focused on language models, reasoning, training, evaluation, and model understanding.
My research takes a data-centric approach to understanding language-model abilities: how training data can shape which capabilities emerge, how data exposure can reveal generalization during pre-training, and how evaluation data can uncover capabilities and failures that final-answer benchmarks miss.
I study how the structure of pre-training data affects which capabilities language models acquire. By simplifying the pre-training distribution, I investigate whether abilities typically associated with larger models—such as zero-shot in-context learning—can emerge at much smaller scales.
I use data exposure during pre-training as a way to study when models move beyond what they have directly seen. By separating exposed and unexposed grammatical examples, I trace delayed generalization across training and analyze how model representations and attention change as this behavior emerges.
I develop process-aware evaluations for reasoning capabilities that are difficult to capture with standard final-answer benchmarks. Using the CROWDMATH program we develop an annotated dataset with collaborative research discussions and study whether language models can recognize partial progress, errors, corrections, and completed results in collaborative mathematical problem solving.
I have also worked on making large-scale language-model training more parameter-efficient. In ReLoRA, we introduce a method for training high-rank neural networks through a sequence of low-rank updates, showing that Transformer language models can achieve performance comparable to standard full-rank training while reducing memory use and improving training speed. The work explores whether techniques commonly used for parameter-efficient adaptation can also make pre-training more efficient.