
Saumya Gandhi
AI/ML Engineer and Researcher passionate about large language models
saumya-gandhi
Bellevue, Washington, United States
Joined February 2025
Network
2.2K connectionsSummary
Saumya Gandhi is an emerging AI/ML engineer and researcher with a strong academic foundation from Carnegie Mellon University and Visvesvaraya National Institute of Technology. She specializes in large language models, natural language processing, and the development of high-quality synthetic data for AI applications. Her recent role at Anthropic as a Member of Technical Staff in the Applied AI Fine-tuning team, following her leadership in model quality at OpenPipe, highlights her dedication to advancing powerful and safe AI. github+2
Her research contributions include developing 'DataTune,' a novel method for generating retrieval-augmented synthetic data, and multiple publications on suicide ideation detection using transformer-based models and deep adversarial learning. This body of work demonstrates her expertise in practical applications of NLP and her commitment to creating robust and high-performing AI systems, particularly with a focus on societal impact in mental healthcare. github+2
Saumya has diverse experience across major tech companies and research institutions, including a Summer Analyst role at Goldman Sachs, a Research Assistant position at Oracle, and various internships focusing on AI and software development. Her early career also includes co-founding a debate club and mentoring students, showcasing leadership and a proactive approach to learning and community involvement. github+1
Work
Education
Projects
Writing
Towards Ordinal Suicide Ideation Detection on Social Media
January 1, 2021Research co-authored and presented at WSDM '2021.
A Time-Aware Transformer Based Model for Suicide Ideation Detection on Social Media
January 1, 2020Research co-authored and presented at EMNLP '2020.
Better Synthetic Data by Retrieving and Transforming Existing Datasets (DataTune)
Introduced DataTune, a method to create retrieval augmented synthetic data grounded in publicly available datasets. This approach autonomously retrieves and transforms data from sources like Hugging Face to meet specific task requirements, offering a high-quality alternative to LLM-generated synthetic data.
Robust Suicide Risk Assessment on Social Media via Deep Adversarial Learning
Research co-authored and published in JAMIA.
Hobbies
Enjoys playing chess. github
Similar profiles
DD
Deedy Das
Partner at Menlo Ventures
110.9K connections
VSVaruni Sarwal
Chief Executive Officer at TriFetch
23K connections
EMEvania Muljono
23.2K connections
DPDevi Parikh
Co-Founder at Yutori
15K connections
ASAdvaith Sridhar
Co-Founder at Discovered Materials
6.1K connections
RSRachitt Shah
AI at Accel
22.7K connections