Skip to content
View snjev310's full-sized avatar
🏠
Working from home
🏠
Working from home

Highlights

  • Pro

Block or report snjev310

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
snjev310/README.md

Hi, I'm Sanjeev Kumar 👋

TCS Research Fellow · PhD Student at IIT Bombay (CSE) Working on NLP for Extremely Low-Resource Languages


About Me

I am a fifth-year PhD student at the Department of Computer Science and Engineering, IIT Bombay, supervised by Prof. Preethi Jyothi and Late Prof. Pushpak Bhattacharyya.

My research focuses on extremely low-resource languages and how to transfer knowledge from large pretrained models to languages that have almost no data.

I was recently a Visiting Research Scholar at the University of Sheffield, hosted by Prof. Nikos Aletras.


Research Interests

  • Machine Translation for Extremely Low-Resource Languages
  • Byte-level and tokenizer-free NLP
  • Morphosyntactic Analysis (POS tagging, NER)
  • Automatic Speech Recognition for Low-Resource Languages
  • Cross-lingual Transfer Learning

Recent Papers

  • [EMNLP 2026] When Tokenizers Fail: Byte-Level Chunking for Zero-Shot Transfer to Low-Resource Languages · Paper · Code

  • [EACL 2026] SrcMix: Mixing of Related Source Languages Benefits Extremely Low-resource Machine Translation · Paper · Code

  • [ACL 2024] Part-of-speech Tagging for Extremely Low-resource Indian Languages · Paper · Code

Pinned Loading

  1. acl-24-pos acl-24-pos Public

    Repository for the ACL 2024 paper "Part-of-Speech Tagging for Extremely Low-resource Indian Languages" (Angika, Magahi, and Bhojpuri). Includes datasets, zero-shot baselines, and look-back inferenc…

    Python 1

  2. SrcMix SrcMix Public

    Repository for the EACL 2026 paper "SrcMix: Mixing of Related Source Languages Benefits Extremely Low-resource Machine Translation". Contains the dataset and code for SrcMix, a controlled source-la…

    Python 1

  3. ByteChunk ByteChunk Public

    Repository for the EMNLP 2026 paper "When Tokenizers Fail: Byte-Level Chunking for Zero-Shot Transfer to Low-Resource Languages".

    Python