Sherin Muckatira

Sherin Muckatira

Research Interests: Language Models, Pre-training, Generalization, LLM Evaluation, Mathematical Reasoning, Interpretability, Parameter-Efficient Fine-Tuning

I’m a Computer Science PhD candidate at the University of Massachusetts Lowell, advised by Prof. Anna Rumshisky. My research takes a data-centric approach to understanding language-model abilities: how training data can shape which capabilities emerge, how data exposure can reveal when models move from memorization toward generalization, and how evaluation data can uncover failures in LLMs that standard benchmarks miss. My work spans language-model pre-training, training dynamics, generalization, model analysis and interpretability, and LLM evaluation. I’m interested in Research Scientist and Applied Scientist roles focused on language models, reasoning, training, evaluation, and model understanding.

News

[Aug 2026] Successfully defended my PhD dissertation proposal.
[Aug 2026] CROWDMATH paper accepted to Findings of EMNLP 2026.
[Aug 2025] Completed my second Applied Scientist internship at Amazon.
[May 2024] Emergent Abilities accepted to Findings of NAACL 2024.
[May 2024] Completed an Applied Science internship at Amazon.
[March 2024] Our work on ReLoRA accepted to ICLR 2024!

Research

My research takes a data-centric approach to understanding language-model abilities: how training data can shape which capabilities emerge, how data exposure can reveal generalization during pre-training, and how evaluation data can uncover capabilities and failures that final-answer benchmarks miss.

Eliciting abilities through pre-training data

I study how the structure of pre-training data affects which capabilities language models acquire. By simplifying the pre-training distribution, I investigate whether abilities typically associated with larger models—such as zero-shot in-context learning—can emerge at much smaller scales.

Understanding generalization during pre-training

I use data exposure during pre-training as a way to study when models move beyond what they have directly seen. By separating exposed and unexposed grammatical examples, I trace delayed generalization across training and analyze how model representations and attention change as this behavior emerges.

Evaluating reasoning beyond final answers

I develop process-aware evaluations for reasoning capabilities that are difficult to capture with standard final-answer benchmarks. Using the CROWDMATH program we develop an annotated dataset with collaborative research discussions and study whether language models can recognize partial progress, errors, corrections, and completed results in collaborative mathematical problem solving.

Efficient language-model pre-training

I have also worked on making large-scale language-model training more parameter-efficient. In ReLoRA, we introduce a method for training high-rank neural networks through a sequence of low-rank updates, showing that Transformer language models can achieve performance comparable to standard full-rank training while reducing memory use and improving training speed. The work explores whether techniques commonly used for parameter-efficient adaptation can also make pre-training more efficient.

Publications

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions
S. Muckatira, J. Geneson, S. Gerovitch, P. Etingof, M. Gronas, A. Rumshisky
EMNLP Findings, 2026
A Pre-Training Analogue of Grokking in Language Models: Tracing Delayed Grammatical Generalization
S. Muckatira, N. Shivagunde, V. Deshpande, A. Rumshisky
Under Review, 2026
Playing with Words, Improving with Rewards: Training Language Models for Creative Association
V. Deshpande, N. Shivagunde, S. Muckatira, H. Glaude, M. Gronas, C. Stevenson, R. Beaty, A. Rumshisky
Under Review, 2026
AGC-Bench: Measuring Artificial General Creativity
R. E. Beaty, V. Deshpande, A. Attuch, P. V. DiStefano, C. K. Y. Lai, S. Muckatira, R. Pujari, S. Roy, N. Shivagunde, C. E. Stevenson, M. Gronas, A. Rumshisky
Under Review, 2026
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
N. Shivagunde, V. Deshpande, S. Muckatira, A. Rumshisky
Under Review, 2026
Emergent Abilities in Reduced-Scale Generative Language Models
S. Muckatira, V. Deshpande, V. Lialin, A. Rumshisky
NAACL Findings, 2024
ReLoRA: High-Rank Training Through Low-Rank Updates
V. Lialin, S. Muckatira, N. Shivagunde, A. Rumshisky
ICLR, 2024
Deconstructing In-Context Learning: Understanding Prompts via Corruption
N. Shivagunde, V. Lialin, S. Muckatira, A. Rumshisky
LREC-Coling, 2024
Let's Reinforce Step by Step
S. Pan, V. Lialin, S. Muckatira, A. Rumshisky
NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following
Properties Of Winning Tickets On Skin Lesion Classification
S. Muckatira
ECCV WiCV Workshop, 2020
Image-level and group-level models for Drosophila gene expression pattern annotation
Q. Sun, S. Muckatira, L. Yuan, S. Ji, S. Newfeld, S. Kumar, J. Ye
BMC bioinformatics, 2013

Education

University of Massachusetts Lowell Lowell, Massachusetts
Ph.D. in Computer Science; GPA 4.00 September 2021 – Present
Advisor: Prof. Anna Rumshisky
Arizona State University Tempe, Arizona
Master of Science in Electrical Engineering; GPA 3.79 August 2011 – May 2013
Sir M Visvesvaraya Institute of Technology Bangalore, India
Bachelor of Engineering in Electronics and Communication; GPA 4.00 September 2007 – July 2011

Experience

Research Assistant Lowell, MA
University of Massachusetts Lowell | May 2023 – Present
PI: Prof. Anna Rumshisky
  • Study how language-model capabilities emerge during pre-training, with a focus on training dynamics, generalization, memorization, and learned representations.
  • Pre-trained and analyzed Transformer language models across multiple scales to investigate when linguistic capabilities emerge and how they relate to data exposure.
  • Developed data-centric evaluation methods for LLM reasoning, including benchmarks for mathematical problem solving and recognizing incremental progress toward solutions.
  • Developing and testing efficient pre-training methods that reduce computational cost without sacrificing performance.
  • Research spans LLM pre-training, evaluation, model interpretability, efficient training, and reasoning; published at NAACL, EMNLP, ICLR, and LREC-COLING.
Applied Scientist Intern Boston, MA
Amazon | May 2025 – August 2025
PI: Dr. Rinat Khaziev
  • Developed multilingual evaluation methods for large language models, studying model quality beyond English.
  • Built and trained LLM judge models for evaluating multi-turn conversations across languages.
Applied Scientist Intern Remote
Amazon | May 2024 – August 2024
PI: Ikkei Itoku
  • Built a synthetic-data generation pipeline for domains with limited and privacy-sensitive training data.
  • Created structured training datasets and fine-tuned Mistral-7B-Instruct for document understanding and guideline identification.
Research Aide Tempe, AZ
Arizona State University | July 2012 – May 2013
PI: Prof. Jieping Ye
  • Implemented Gene expression pattern annotation using SIFT feature extraction on images in the Berkeley Drosophila Genome Project (BDGP).
  • Constructed Codebooks using Bag of Words and Sparse Coding Approach.
Senior Software Engineer Boxborough, MA
Qualcomm | October 2016 – December 2021
  • Developed firmware for the physical layer of Wireless LAN chips using the Wifi 802.11 protocol.
  • Designed and implemented features such as Spectral Scan and Radar Detection.
Applications Software Engineer Chandler, AZ
NXP | June 2013 – October 2016
  • Developed signal processing applications for radio communication, focusing on transmit and receive chains on a Vector Signal Processor for Power Amplifier characterization.
  • Implemented communication interfaces between host processors and co-processors to enhance functionality in Power Amplifier characterization applications.