Center for Tamil Natural Language Processing Research

Center for Tamil Natural Language Processing Research Contact information, map and directions, contact form, opening hours, services, ratings, photos, videos and announcements from Center for Tamil Natural Language Processing Research, Education, 63, Sir Pon, Thirunelvelly, Ramanathan Road, Kallady, Jaffna.

Center for Tamil natural language processing research aims to research and develop natural language processing tools required for Tamil and to build an active scholarly network of people contributing to the advancement of the language. The Center for Tamil natural language processing research aims to research and develop natural language processing tools required for Tamil and to build an active scholarly network of people contributing to the advancement of the language.

๐Ÿš€ ๐—•๐˜‚๐—ถ๐—น๐—ฑ๐—ถ๐—ป๐—ด ๐—ฎ ๐—™๐—ถ๐—ป๐—ฒ-๐—š๐—ฟ๐—ฎ๐—ถ๐—ป๐—ฒ๐—ฑ ๐—ง๐—ฎ๐—บ๐—ถ๐—น ๐—ก๐—ฎ๐—บ๐—ฒ๐—ฑ ๐—˜๐—ป๐˜๐—ถ๐˜๐˜† ๐—ฅ๐—ฒ๐—ฐ๐—ผ๐—ด๐—ป๐—ถ๐˜๐—ถ๐—ผ๐—ป (๐—™๐—ด๐—ก๐—˜๐—ฅ) ๐—ฆ๐˜†๐˜€๐˜๐—ฒ๐—บ ๐˜„๐—ถ๐˜๐—ต ๐—š๐—ผ๐—ผ๐—ด๐—น๐—ฒ ๐— ๐˜‚๐—ฅ๐—œ๐—ŸNamed Entity Recognition (๐—ก๐—˜๐—ฅ) ...
10/07/2026

๐Ÿš€ ๐—•๐˜‚๐—ถ๐—น๐—ฑ๐—ถ๐—ป๐—ด ๐—ฎ ๐—™๐—ถ๐—ป๐—ฒ-๐—š๐—ฟ๐—ฎ๐—ถ๐—ป๐—ฒ๐—ฑ ๐—ง๐—ฎ๐—บ๐—ถ๐—น ๐—ก๐—ฎ๐—บ๐—ฒ๐—ฑ ๐—˜๐—ป๐˜๐—ถ๐˜๐˜† ๐—ฅ๐—ฒ๐—ฐ๐—ผ๐—ด๐—ป๐—ถ๐˜๐—ถ๐—ผ๐—ป (๐—™๐—ด๐—ก๐—˜๐—ฅ) ๐—ฆ๐˜†๐˜€๐˜๐—ฒ๐—บ ๐˜„๐—ถ๐˜๐—ต ๐—š๐—ผ๐—ผ๐—ด๐—น๐—ฒ ๐— ๐˜‚๐—ฅ๐—œ๐—Ÿ

Named Entity Recognition (๐—ก๐—˜๐—ฅ) is one of the core tasks in ๐—ก๐—ฎ๐˜๐˜‚๐—ฟ๐—ฎ๐—น ๐—Ÿ๐—ฎ๐—ป๐—ด๐˜‚๐—ฎ๐—ด๐—ฒ ๐—ฃ๐—ฟ๐—ผ๐—ฐ๐—ฒ๐˜€๐˜€๐—ถ๐—ป๐—ด (๐—ก๐—Ÿ๐—ฃ), enabling AI systems to identify and classify real-world entities such as people, organizations, locations, products, events, and more.

Most existing Tamil NER systems focus only on three coarse-grained categoriesโ€”๐—ฃ๐—˜๐—ฅ๐—ฆ๐—ข๐—ก, ๐—Ÿ๐—ข๐—–๐—”๐—ง๐—œ๐—ข๐—ก, and ๐—ข๐—ฅ๐—š๐—”๐—ก๐—œ๐—ญ๐—”๐—ง๐—œ๐—ข๐—ก. While suitable for basic information extraction, these categories are often insufficient for advanced document intelligence, semantic search, and knowledge graph applications.

As part of ongoing research at ๐—–๐—ง๐—ก๐—Ÿ๐—ฃ๐—ฅ, we explored a different direction by fine-tuning ๐—š๐—ผ๐—ผ๐—ด๐—น๐—ฒ ๐— ๐˜‚๐—ฅ๐—œ๐—Ÿ on the ๐—ฆ๐—ฎ๐—บ๐—ฝ๐˜‚๐—ฟ๐—ก๐—˜๐—ฅ Fine-Grained Tamil NER dataset to enable richer semantic understanding of Tamil documents.

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—ช๐—ต๐˜† ๐—™๐—ถ๐—ป๐—ฒ-๐—š๐—ฟ๐—ฎ๐—ถ๐—ป๐—ฒ๐—ฑ ๐—ก๐—˜๐—ฅ?
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

Instead of assigning broad labels like PERSON or ORGANIZATION, Fine-Grained NER recognizes much more specific entity types.

Examples include:

โœ“ Government Organizations
โœ“ Educational Institutions
โœ“ Hospitals
โœ“ Buildings
โœ“ Products
โœ“ Events
โœ“ Diseases
โœ“ Creative Works
โœ“ Athletes
โœ“ Artists
โœ“ Writers
โœ“ Scholars
โœ“ Biological Entities

This richer semantic representation significantly improves downstream AI systems such as ๐—œ๐—ป๐—ณ๐—ผ๐—ฟ๐—บ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—˜๐˜…๐˜๐—ฟ๐—ฎ๐—ฐ๐˜๐—ถ๐—ผ๐—ป, ๐—ž๐—ป๐—ผ๐˜„๐—น๐—ฒ๐—ฑ๐—ด๐—ฒ ๐—š๐—ฟ๐—ฎ๐—ฝ๐—ต๐˜€, ๐—ฆ๐—ฒ๐—บ๐—ฎ๐—ป๐˜๐—ถ๐—ฐ ๐—ฆ๐—ฒ๐—ฎ๐—ฟ๐—ฐ๐—ต, and ๐—ฅ๐—ฒ๐˜๐—ฟ๐—ถ๐—ฒ๐˜ƒ๐—ฎ๐—น-๐—”๐˜‚๐—ด๐—บ๐—ฒ๐—ป๐˜๐—ฒ๐—ฑ ๐—š๐—ฒ๐—ป๐—ฒ๐—ฟ๐—ฎ๐˜๐—ถ๐—ผ๐—ป (๐—ฅ๐—”๐—š).

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—ช๐—ต๐˜† ๐— ๐˜‚๐—ฅ๐—œ๐—Ÿ?
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

Unlike general multilingual transformer models, ๐— ๐˜‚๐—ฅ๐—œ๐—Ÿ was specifically pretrained for Indian languages using:

โœ“ Large-scale Indian language corpora
โœ“ Parallel translated sentence pairs
โœ“ Transliterated document pairs
โœ“ Multilingual Masked Language Modeling

Because Tamil is one of MuRIL's primary target languages, it provides a strong contextual representation for Fine-Grained Tamil Named Entity Recognition.

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐— ๐—ผ๐—ฑ๐—ฒ๐—น ๐—”๐—ฟ๐—ฐ๐—ต๐—ถ๐˜๐—ฒ๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

The overall system follows a transformer-based token classification pipeline:

๐—ฆ๐—ฎ๐—บ๐—ฝ๐˜‚๐—ฟ๐—ก๐—˜๐—ฅ ๐——๐—ฎ๐˜๐—ฎ๐˜€๐—ฒ๐˜
โ†“
MuRIL Tokenizer
โ†“
WordPiece Tokenization
โ†“
BIO Sequence Labeling
โ†“
Transformer-Based Token Classification
โ†“
Fine-Grained Entity Prediction

Rather than modifying the transformer architecture, the model specializes MuRIL through supervised fine-tuning for fine-grained semantic classification.

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—ง๐—ฒ๐—ฐ๐—ต๐—ป๐—ถ๐—ฐ๐—ฎ๐—น ๐—–๐—ต๐—ฎ๐—น๐—น๐—ฒ๐—ป๐—ด๐—ฒ๐˜€
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

During experimentation, several challenges were observed:

โœ“ High-cardinality semantic label space
โœ“ Agglutinative Tamil morphology
โœ“ Long multi-token entity spans
โœ“ Fine-grained semantic ambiguity
โœ“ Context-dependent classification

These challenges require the model to rely on contextual semantic understanding rather than simple lexical matching.

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐— ๐—ผ๐—ฑ๐—ฒ๐—น ๐—ข๐—ฏ๐˜€๐—ฒ๐—ฟ๐˜ƒ๐—ฎ๐˜๐—ถ๐—ผ๐—ป๐˜€
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

The fine-tuned MuRIL model demonstrated strong performance across multiple entity categories, including:

โœ“ Person
โœ“ Organization
โœ“ Government Organization
โœ“ Educational Institution
โœ“ Buildings
โœ“ Hospitals
โœ“ Products
โœ“ Creative Works

The multilingual representations learned during MuRIL pretraining transferred effectively to Fine-Grained Tamil NER through supervised task adaptation.

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—”๐—ฝ๐—ฝ๐—น๐—ถ๐—ฐ๐—ฎ๐˜๐—ถ๐—ผ๐—ป๐˜€
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

Fine-Grained Tamil NER enables richer semantic understanding for:

โœ“ Information Extraction
โœ“ Entity Linking
โœ“ Semantic Search
โœ“ Knowledge Graph Population
โœ“ Retrieval-Augmented Generation (RAG)
โœ“ Digital Libraries
โœ“ Document Intelligence
โœ“ Tamil Question Answering
โœ“ Large-Scale Archive Processing

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—ž๐—ฒ๐˜† ๐—ง๐—ฒ๐—ฐ๐—ต๐—ป๐—ถ๐—ฐ๐—ฎ๐—น ๐—œ๐—ป๐˜€๐—ถ๐—ด๐—ต๐˜
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

One of the key findings is that the limitation of conventional Tamil NER is not entity detection itself, but semantic granularity.

Identifying an entity simply as PERSON or ORGANIZATION provides limited value for downstream reasoning. Fine-Grained NER preserves domain-specific semantic distinctions, enabling richer knowledge representation and more expressive AI applications.

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—œ๐—บ๐—ฝ๐—น๐—ฒ๐—บ๐—ฒ๐—ป๐˜๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ฎ๐˜ ๐—–๐—ง๐—ก๐—Ÿ๐—ฃ๐—ฅ
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

The current implementation provides:

โœ“ Transformer-based token classification using MuRIL
โœ“ BIO sequence labeling
โœ“ Fine-grained semantic entity classification
โœ“ High-cardinality entity taxonomy
โœ“ Context-aware multilingual representations
โœ“ Integration-ready outputs for Information Extraction and Knowledge Graph Construction

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—™๐˜‚๐—น๐—น ๐—ง๐—ฒ๐—ฐ๐—ต๐—ป๐—ถ๐—ฐ๐—ฎ๐—น ๐—•๐—น๐—ผ๐—ด
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

๐Ÿ”— https://www.ctnlpr.com/2026/07/08/building-a-fine-grained-tamil-named-entity-recognition-system-with-muril/

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

๐Ÿš€ ๐—•๐˜‚๐—ถ๐—น๐—ฑ๐—ถ๐—ป๐—ด ๐—ฎ ๐—ง๐—ฟ๐—ฎ๐—ป๐˜€๐—ณ๐—ผ๐—ฟ๐—บ๐—ฒ๐—ฟ-๐—•๐—ฎ๐˜€๐—ฒ๐—ฑ ๐—ง๐—ฎ๐—บ๐—ถ๐—น ๐—ฅ๐—ฒ๐—น๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—˜๐˜…๐˜๐—ฟ๐—ฎ๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐—ฃ๐—ถ๐—ฝ๐—ฒ๐—น๐—ถ๐—ป๐—ฒ ๐—ณ๐—ผ๐—ฟ ๐—ž๐—ป๐—ผ๐˜„๐—น๐—ฒ๐—ฑ๐—ด๐—ฒ ๐—š๐—ฟ๐—ฎ๐—ฝ๐—ต ๐—–๐—ผ๐—ป๐˜€๐˜๐—ฟ๐˜‚๐—ฐ๐˜๐—ถ๐—ผ๐—ปKnowledge Graph constr...
24/06/2026

๐Ÿš€ ๐—•๐˜‚๐—ถ๐—น๐—ฑ๐—ถ๐—ป๐—ด ๐—ฎ ๐—ง๐—ฟ๐—ฎ๐—ป๐˜€๐—ณ๐—ผ๐—ฟ๐—บ๐—ฒ๐—ฟ-๐—•๐—ฎ๐˜€๐—ฒ๐—ฑ ๐—ง๐—ฎ๐—บ๐—ถ๐—น ๐—ฅ๐—ฒ๐—น๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—˜๐˜…๐˜๐—ฟ๐—ฎ๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐—ฃ๐—ถ๐—ฝ๐—ฒ๐—น๐—ถ๐—ป๐—ฒ ๐—ณ๐—ผ๐—ฟ ๐—ž๐—ป๐—ผ๐˜„๐—น๐—ฒ๐—ฑ๐—ด๐—ฒ ๐—š๐—ฟ๐—ฎ๐—ฝ๐—ต ๐—–๐—ผ๐—ป๐˜€๐˜๐—ฟ๐˜‚๐—ฐ๐˜๐—ถ๐—ผ๐—ป

Knowledge Graph construction from unstructured text remains one of the fundamental challenges in ๐—ก๐—ฎ๐˜๐˜‚๐—ฟ๐—ฎ๐—น ๐—Ÿ๐—ฎ๐—ป๐—ด๐˜‚๐—ฎ๐—ด๐—ฒ ๐—ฃ๐—ฟ๐—ผ๐—ฐ๐—ฒ๐˜€๐˜€๐—ถ๐—ป๐—ด (๐—ก๐—Ÿ๐—ฃ).

For low-resource languages such as ๐—ฆ๐—ฟ๐—ถ ๐—Ÿ๐—ฎ๐—ป๐—ธ๐—ฎ๐—ป ๐—ง๐—ฎ๐—บ๐—ถ๐—น, this challenge is even more significant due to limited annotated datasets, domain-specific resources, and production-ready information extraction systems.

As part of ongoing research at ๐—–๐—ง๐—ก๐—Ÿ๐—ฃ๐—ฅ, we designed and implemented a ๐—ง๐—ฟ๐—ฎ๐—ป๐˜€๐—ณ๐—ผ๐—ฟ๐—บ๐—ฒ๐—ฟ-๐—ฏ๐—ฎ๐˜€๐—ฒ๐—ฑ ๐—ง๐—ฎ๐—บ๐—ถ๐—น ๐—ฅ๐—ฒ๐—น๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—˜๐˜…๐˜๐—ฟ๐—ฎ๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐—ฝ๐—ถ๐—ฝ๐—ฒ๐—น๐—ถ๐—ป๐—ฒ capable of converting unstructured Tamil documents into structured knowledge triples.

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—ฅ๐—ฒ๐˜€๐—ฒ๐—ฎ๐—ฟ๐—ฐ๐—ต ๐— ๐—ผ๐˜๐—ถ๐˜ƒ๐—ฎ๐˜๐—ถ๐—ผ๐—ป
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

Consider the Tamil sentence:

เฎšเฎพเฎฎเฏ.เฎ.เฎšเฎชเฎพเฎชเฎคเฎฟ เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃเฎคเฏเฎคเฎฟเฎฒเฏ เฎตเฎพเฎดเฏเฎจเฏเฎคเฎพเฎฐเฏ.

Humans can easily infer:

(เฎšเฎพเฎฎเฏ.เฎ.เฎšเฎชเฎพเฎชเฎคเฎฟ, LIVED_IN, เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃเฎฎเฏ)

However, an NLP system must solve multiple sub-problems before arriving at this representation:

โœ“ Sentence Segmentation
โœ“ Named Entity Recognition
โœ“ Entity Typing
โœ“ Entity Pair Generation
โœ“ Relation Classification
โœ“ Structured Triple Construction

These tasks form the foundation of modern ๐—œ๐—ป๐—ณ๐—ผ๐—ฟ๐—บ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—˜๐˜…๐˜๐—ฟ๐—ฎ๐—ฐ๐˜๐—ถ๐—ผ๐—ป and ๐—ž๐—ป๐—ผ๐˜„๐—น๐—ฒ๐—ฑ๐—ด๐—ฒ ๐—š๐—ฟ๐—ฎ๐—ฝ๐—ต systems.

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—ฃ๐—ถ๐—ฝ๐—ฒ๐—น๐—ถ๐—ป๐—ฒ ๐—”๐—ฟ๐—ฐ๐—ต๐—ถ๐˜๐—ฒ๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

The system follows a modular pipeline design:

๐——๐—ผ๐—ฐ๐˜‚๐—บ๐—ฒ๐—ป๐˜
โ†’ Sentence Splitter
โ†’ Named Entity Recognition
โ†’ Pair Generator
โ†’ Entity Marker
โ†’ Relation Extraction
โ†’ Triple Builder

Unlike end-to-end extraction architectures, a pipeline-based approach was selected due to:

โœ“ Modularity
โœ“ Interpretability
โœ“ Ease of debugging
โœ“ Suitability for low-resource NLP environments

Each component can be independently evaluated, improved, or replaced without affecting the overall system.

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—–๐˜‚๐—ฟ๐—ฟ๐—ฒ๐—ป๐˜ ๐—–๐—ฎ๐—ฝ๐—ฎ๐—ฏ๐—ถ๐—น๐—ถ๐˜๐—ถ๐—ฒ๐˜€
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

The current implementation supports:

โœ“ Sentence Segmentation

โœ“ Named Entity Recognition

Entity Types:

โ€ข PERSON
โ€ข LOCATION
โ€ข ORGANIZATION

โœ“ Entity Pair Generation

Supported Semantic Pair Constraints:

โ€ข PERSON โ†” LOCATION
โ€ข PERSON โ†” ORGANIZATION
โ€ข ORGANIZATION โ†” LOCATION

โœ“ Transformer-Based Relation Extraction

Current Relation Labels:

โ€ข LIVED_IN
โ€ข WORKED_AT
โ€ข CONTRIBUTED_TO
โ€ข PART_OF
โ€ข BORN_IN

โœ“ Knowledge Triple Construction

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—ง๐—ฒ๐—ฐ๐—ต๐—ป๐—ถ๐—ฐ๐—ฎ๐—น ๐—”๐—ฝ๐—ฝ๐—ฟ๐—ผ๐—ฎ๐—ฐ๐—ต
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

The current system uses a multilingual ๐—ก๐—ฎ๐˜๐˜‚๐—ฟ๐—ฎ๐—น ๐—Ÿ๐—ฎ๐—ป๐—ด๐˜‚๐—ฎ๐—ด๐—ฒ ๐—œ๐—ป๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ (๐—ก๐—Ÿ๐—œ) Transformer model in a ๐˜‡๐—ฒ๐—ฟ๐—ผ-๐˜€๐—ต๐—ผ๐˜ setting for relation classification.

This approach enables rapid experimentation without requiring a dedicated Tamil Relation Extraction dataset.

Entity mentions are explicitly marked before classification:

[E1] เฎšเฎพเฎฎเฏเฎชเฎšเฎฟเฎตเฎฎเฏ [/E1]

[E2] เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃเฎคเฏเฎคเฎฟเฎฒเฏ [/E2]

This helps the transformer focus on target entities when predicting semantic relationships.

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—˜๐—ป๐—ด๐—ถ๐—ป๐—ฒ๐—ฒ๐—ฟ๐—ถ๐—ป๐—ด ๐—–๐—ต๐—ฎ๐—น๐—น๐—ฒ๐—ป๐—ด๐—ฒ๐˜€
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

During experimentation, several practical challenges were observed:

โœ“ Tamil Unicode combining-character boundary issues

โœ“ Tokenization inconsistencies

โœ“ Entity span normalization

โœ“ Candidate relation filtering

โœ“ Low-resource language constraints

To address these issues, an entity post-processing layer was introduced to improve boundary consistency and downstream extraction quality.

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—˜๐˜…๐—ฝ๐—ฒ๐—ฟ๐—ถ๐—บ๐—ฒ๐—ป๐˜๐—ฎ๐—น ๐—ข๐˜‚๐˜๐—ฐ๐—ผ๐—บ๐—ฒ๐˜€
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

The pipeline was evaluated on:

โœ“ Personal biographies

โœ“ Academic affiliations

โœ“ Institutional relationships

โœ“ Historical narratives

โœ“ Sri Lankan Tamil named entities

The experiments demonstrate that the architecture can successfully extract meaningful knowledge triples from Tamil text and provide a strong foundation for future Knowledge Graph construction.

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—™๐˜‚๐˜๐˜‚๐—ฟ๐—ฒ ๐—ฅ๐—ฒ๐˜€๐—ฒ๐—ฎ๐—ฟ๐—ฐ๐—ต ๐——๐—ถ๐—ฟ๐—ฒ๐—ฐ๐˜๐—ถ๐—ผ๐—ป๐˜€
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

Coreference Resolution
โ†“
Entity Canonicalization
โ†“
Ontology Validation
โ†“
Relation Filtering
โ†“
Knowledge Graph Construction

Additional research areas include:

โœ“ MuRIL-based NER

โœ“ XLM-R-based NER

โœ“ Fine-tuned Tamil Relation Extraction Models

โœ“ Document-Level Relation Extraction

โœ“ Ontology-Aware Triple Validation

โœ“ Knowledge Graph Population from Noolaham Archives

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”
๐—™๐˜‚๐—น๐—น ๐—ง๐—ฒ๐—ฐ๐—ต๐—ป๐—ถ๐—ฐ๐—ฎ๐—น ๐—•๐—น๐—ผ๐—ด
https://www.ctnlpr.com/2026/06/23/building-a-transformer-based-tamil-relation-extraction-pipeline-for-knowledge-graph-construction/
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

๐Ÿš€๐—•๐˜‚๐—ถ๐—น๐—ฑ๐—ถ๐—ป๐—ด ๐—ฎ ๐—ง๐—ฎ๐—บ๐—ถ๐—น ๐—–๐—ผ๐—ฟ๐—ฒ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—ฅ๐—ฒ๐˜€๐—ผ๐—น๐˜‚๐˜๐—ถ๐—ผ๐—ป ๐—ฆ๐˜†๐˜€๐˜๐—ฒ๐—บ ๐—ณ๐—ผ๐—ฟ ๐—œ๐—ป๐—ณ๐—ผ๐—ฟ๐—บ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—˜๐˜…๐˜๐—ฟ๐—ฎ๐—ฐ๐˜๐—ถ๐—ผ๐—ปModern Information Extraction systems rely on ...
15/06/2026

๐Ÿš€๐—•๐˜‚๐—ถ๐—น๐—ฑ๐—ถ๐—ป๐—ด ๐—ฎ ๐—ง๐—ฎ๐—บ๐—ถ๐—น ๐—–๐—ผ๐—ฟ๐—ฒ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—ฅ๐—ฒ๐˜€๐—ผ๐—น๐˜‚๐˜๐—ถ๐—ผ๐—ป ๐—ฆ๐˜†๐˜€๐˜๐—ฒ๐—บ ๐—ณ๐—ผ๐—ฟ ๐—œ๐—ป๐—ณ๐—ผ๐—ฟ๐—บ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—˜๐˜…๐˜๐—ฟ๐—ฎ๐—ฐ๐˜๐—ถ๐—ผ๐—ป

Modern Information Extraction systems rely on more than Named Entity Recognition (NER). While NER can identify entities such as people, locations, and organizations, it does not explain how references to those entities evolve throughout a document.

๐—ง๐—ต๐—ถ๐˜€ ๐—ถ๐˜€ ๐˜„๐—ต๐—ฒ๐—ฟ๐—ฒ ๐—–๐—ผ๐—ฟ๐—ฒ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—ฅ๐—ฒ๐˜€๐—ผ๐—น๐˜‚๐˜๐—ถ๐—ผ๐—ป (๐—–๐—ฅ) ๐—ฏ๐—ฒ๐—ฐ๐—ผ๐—บ๐—ฒ๐˜€ ๐—ฒ๐˜€๐˜€๐—ฒ๐—ป๐˜๐—ถ๐—ฎ๐—น.

Coreference Resolution is the task of determining whether multiple mentions within a document refer to the same real-world entity. It acts as a critical bridge between entity recognition and structured knowledge extraction.

๐—”๐—ฝ๐—ฝ๐—น๐—ถ๐—ฐ๐—ฎ๐˜๐—ถ๐—ผ๐—ป๐˜€:

โ€ข Knowledge Graph Construction
โ€ข Relation Extraction
โ€ข Semantic Search
โ€ข Document Intelligence
โ€ข Retrieval-Augmented Generation (RAG)
โ€ข Conversational AI

Accurate coreference resolution is often the difference between fragmented information and coherent knowledge.

๐—ง๐—ต๐—ฒ ๐—œ๐—ป๐—ณ๐—ผ๐—ฟ๐—บ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—˜๐˜…๐˜๐—ฟ๐—ฎ๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐—ฃ๐—ฟ๐—ผ๐—ฏ๐—น๐—ฒ๐—บ

During our work at CTNLPR, we observed a common challenge in Tamil document processing.

Documents rarely repeat the full entity name in every sentence. Instead, they rely on:

โ€ข Pronouns
โ€ข Possessive references
โ€ข Descriptive noun phrases
โ€ข Location references

Humans resolve these references naturally using context. Machines do not.

When we use coreference resolution The extracted knowledge becomes meaningful and directly usable within downstream systems.

๐—ช๐—ต๐˜† ๐—ช๐—ฒ ๐——๐—ถ๐—ฑ ๐—ก๐—ผ๐˜ ๐—ฆ๐˜๐—ฎ๐—ฟ๐˜ ๐˜„๐—ถ๐˜๐—ต ๐—ก๐—ฒ๐˜‚๐—ฟ๐—ฎ๐—น ๐—–๐—ผ๐—ฟ๐—ฒ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐— ๐—ผ๐—ฑ๐—ฒ๐—น๐˜€

Most modern coreference systems rely on:

โ€ข Transformer-based architectures
โ€ข Mention-ranking models
โ€ข End-to-end neural systems
โ€ข Span-ranking approaches

While highly effective for English and other high-resource languages, they typically require:

โ€ข Large annotated datasets
โ€ข Extensive model training
โ€ข Significant computational resources
โ€ข Language-specific supervision

Tamil currently lacks large-scale publicly available coreference corpora.

Instead of waiting for benchmark datasets, we explored a different direction:

๐—•๐˜‚๐—ถ๐—น๐—ฑ ๐—ฎ ๐—ฑ๐—ฒ๐˜๐—ฒ๐—ฟ๐—บ๐—ถ๐—ป๐—ถ๐˜€๐˜๐—ถ๐—ฐ, ๐—ฒ๐˜…๐—ฝ๐—น๐—ฎ๐—ถ๐—ป๐—ฎ๐—ฏ๐—น๐—ฒ, ๐—ฎ๐—ป๐—ฑ ๐˜๐—ฎ๐˜€๐—ธ-๐—ผ๐—ฟ๐—ถ๐—ฒ๐—ป๐˜๐—ฒ๐—ฑ ๐—ฐ๐—ผ๐—ฟ๐—ฒ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—ณ๐—ฟ๐—ฎ๐—บ๐—ฒ๐˜„๐—ผ๐—ฟ๐—ธ ๐—ผ๐—ฝ๐˜๐—ถ๐—บ๐—ถ๐˜‡๐—ฒ๐—ฑ ๐—ณ๐—ผ๐—ฟ ๐—œ๐—ป๐—ณ๐—ผ๐—ฟ๐—บ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—˜๐˜…๐˜๐—ฟ๐—ฎ๐—ฐ๐˜๐—ถ๐—ผ๐—ป.

Our goal was not to compete with neural benchmarks.

Our goal was to improve relation extraction quality in real-world Tamil document processing.

๐—ฆ๐˜†๐˜€๐˜๐—ฒ๐—บ ๐—”๐—ฟ๐—ฐ๐—ต๐—ถ๐˜๐—ฒ๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ

Tamil Document
โ†“
Text Normalization
โ†“
Sentence Segmentation
โ†“
Mention Detection
โ†“
Entity Normalization
โ†“
Entity Memory
โ†“
Coreference Resolution
โ†“
Coreference Chain Construction
โ†“
Visualization Layer

Each layer contributes toward discourse-level entity understanding.

๐— ๐—ฒ๐—ป๐˜๐—ถ๐—ผ๐—ป ๐——๐—ฒ๐˜๐—ฒ๐—ฐ๐˜๐—ถ๐—ผ๐—ป

The mention detection layer combines multiple strategies:

โ€ข Named Entity Recognition (PERSON, LOCATION, ORGANIZATION)
โ€ข Pronoun Detection
โ€ข Location Reference Detection
โ€ข Rule-Based Noun Phrase Detection

These mentions become candidates for resolution.

๐—˜๐—ป๐˜๐—ถ๐˜๐˜† ๐—ก๐—ผ๐—ฟ๐—บ๐—ฎ๐—น๐—ถ๐˜‡๐—ฎ๐˜๐—ถ๐—ผ๐—ป

Tamil's rich morphology creates multiple surface forms for the same entity.

๐—˜๐˜…๐—ฎ๐—บ๐—ฝ๐—น๐—ฒ:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃเฎฎเฏ

เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃเฎคเฏเฎคเฎฟเฎฒเฏ

เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃเฎคเฏเฎคเฎฟเฎฉเฏ

เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃเฎคเฏเฎคเฎฟเฎฑเฏเฎ•เฏ
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Using Stanza lemmatization:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃเฎคเฏเฎคเฎฟเฎฒเฏ
โ†“
เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃเฎฎเฏ
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

This reduces entity fragmentation and improves linking consistency.

๐—ง๐—ต๐—ฒ ๐—˜๐—ป๐˜๐—ถ๐˜๐˜† ๐— ๐—ฒ๐—บ๐—ผ๐—ฟ๐˜† ๐—Ÿ๐—ฎ๐˜†๐—ฒ๐—ฟ

One of the key design decisions was introducing a lightweight discourse memory.

Instead of neural antecedent scoring, the system maintains contextual entity state:

last_person
last_location
last_org

Whenever a new entity is detected, the corresponding memory state is updated.

This memory acts as the document's discourse context.

๐—–๐—ผ๐—ฟ๐—ฒ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—ฅ๐—ฒ๐˜€๐—ผ๐—น๐˜‚๐˜๐—ถ๐—ผ๐—ป

Once discourse memory has been established, the resolver links newly encountered mentions to previously observed canonical entities.

The system performs:

โ€ข Person Pronoun Resolution
โ€ข Possessive Resolution
โ€ข Location Resolution
โ€ข Rule-Based Noun Phrase Resolution

By maintaining discourse state across sentence boundaries, fragmented references are transformed into consistent entity representations.

This significantly improves downstream relation extraction quality.

๐—–๐—ผ๐—ฟ๐—ฒ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—–๐—ต๐—ฎ๐—ถ๐—ป ๐—–๐—ผ๐—ป๐˜€๐˜๐—ฟ๐˜‚๐—ฐ๐˜๐—ถ๐—ผ๐—ป

Rather than replacing mentions individually, the system groups related mentions into entity clusters.

consider this example:-

เฎšเฎพเฎฎเฏ.เฎ.เฎšเฎชเฎพเฎชเฎคเฎฟ เฎ’เฎฐเฏ เฎชเฎฟเฎฐเฎชเฎฒ เฎšเฎฎเฏ‚เฎ• เฎšเฏ‡เฎตเฎ•เฎฐเฏ. เฎ…เฎตเฎฐเฏ เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃเฎคเฏเฎคเฎฟเฎฒเฏ เฎชเฎฒ เฎ•เฎฒเฏเฎตเฎฟเฎคเฏ เฎคเฎฟเฎŸเฏเฎŸเฎ™เฏเฎ•เฎณเฏˆ เฎฎเฏเฎฉเฏเฎฉเฏ†เฎŸเฏเฎคเฏเฎคเฎพเฎฐเฏ. เฎ‡เฎจเฏเฎค เฎšเฎฎเฏ‚เฎ•เฎšเฏ‡เฎตเฎ•เฎฐเฏ เฎชเฎฒ เฎตเฎฟเฎฐเฏเฎคเฏเฎ•เฎณเฏˆ เฎชเฏ†เฎฑเฏเฎฑเฏเฎณเฏเฎณเฎพเฎฐเฏ. เฎ…เฎตเฎฐเฏ เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃเฎคเฏเฎคเฎฟเฎฒเฏ เฎชเฎฟเฎฑเฎจเฏเฎคเฎพเฎฐเฏ. เฎ…เฎ™เฏเฎ•เฏ เฎ…เฎตเฎฐเฏเฎ•เฏเฎ•เฏ เฎชเฏ†เฎฐเฏเฎฎเฏ เฎฎเฎคเฎฟเฎชเฏเฎชเฏ เฎ‡เฎฐเฏเฎจเฏเฎคเฎคเฏ. เฎ‡เฎตเฎฐเฏ เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃ เฎจเฏ‚เฎฒเฎ•เฎคเฏเฎคเฎฟเฎฉเฏ เฎ‰เฎฐเฏเฎตเฎพเฎ•เฏเฎ•เฎคเฏเฎคเฎฟเฎฒเฏ เฎฎเฏเฎ•เฏเฎ•เฎฟเฎฏ เฎชเฎ™เฏเฎ•เฎพเฎฑเฏเฎฑเฎฟเฎฉเฎพเฎฐเฏ.

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
Entity: เฎšเฎพเฎฎเฏ.เฎ.เฎšเฎชเฎพเฎชเฎคเฎฟ

โ”œโ”€โ”€ เฎšเฎพเฎฎเฏ.เฎ.เฎšเฎชเฎพเฎชเฎคเฎฟ
โ”œโ”€โ”€ เฎ…เฎตเฎฐเฏ
โ”œโ”€โ”€ เฎ‡เฎจเฏเฎค เฎšเฎฎเฏ‚เฎ•เฎšเฏ‡เฎตเฎ•เฎฐเฏ
โ”œโ”€โ”€ เฎ…เฎตเฎฐเฏ
โ””โ”€โ”€ เฎ‡เฎตเฎฐเฏ

Entity: เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃเฎฎเฏ

โ”œโ”€โ”€ เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃเฎฎเฏ
โ””โ”€โ”€ เฎ…เฎ™เฏเฎ•เฏ
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

These chains provide a document-level view of entity references.

Useful for:

โ€ข Debugging
โ€ข Evaluation
โ€ข Knowledge Graph Construction
โ€ข Relation Extraction

๐—ฉ๐—ถ๐˜€๐˜‚๐—ฎ๐—น๐—ถ๐˜‡๐—ฎ๐˜๐—ถ๐—ผ๐—ป & ๐—”๐—ป๐—ฎ๐—น๐˜†๐˜€๐—ถ๐˜€

To support experimentation and validation, we developed a Streamlit-based visualization layer.

Users can:

โ€ข Submit Tamil documents
โ€ข Inspect generated coreference chains
โ€ข Analyze entity clusters
โ€ข Validate resolution decisions

This provides transparency into the resolution process and helps identify weaknesses in rule design.

๐—ž๐—ฒ๐˜† ๐—œ๐—ป๐˜€๐—ถ๐—ด๐—ต๐˜

๐—–๐—ผ๐—ฟ๐—ฒ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—ฅ๐—ฒ๐˜€๐—ผ๐—น๐˜‚๐˜๐—ถ๐—ผ๐—ป ๐—ถ๐˜€ ๐—ป๐—ผ๐˜ ๐—บ๐—ฒ๐—ฟ๐—ฒ๐—น๐˜† ๐—ฎ ๐—ฝ๐—ฟ๐—ผ๐—ป๐—ผ๐˜‚๐—ป-๐—ฟ๐—ฒ๐˜€๐—ผ๐—น๐˜‚๐˜๐—ถ๐—ผ๐—ป ๐˜๐—ฎ๐˜€๐—ธ.

It is an entity consistency layer that connects:

โ€ข Named Entity Recognition
โ€ข Relation Extraction
โ€ข Knowledge Graph Construction
โ€ข Semantic Search
โ€ข RAG Systems

Without coreference:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
(เฎ‡เฎตเฎฐเฏ, เฎชเฎ™เฏเฎ•เฎพเฎฑเฏเฎฑเฎฟเฎฉเฎพเฎฐเฏ
, เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃ เฎจเฏ‚เฎฒเฎ•เฎฎเฏ)
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

With coreference:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
(เฎšเฎพเฎฎเฏ.เฎ.เฎšเฎชเฎพเฎชเฎคเฎฟ, เฎชเฎ™เฏเฎ•เฎพเฎฑเฏเฎฑเฎฟเฎฉเฎพเฎฐเฏ
, เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃ เฎจเฏ‚เฎฒเฎ•เฎฎเฏ)
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

The second representation is immediately usable within structured knowledge systems.

๐—ฆ๐˜†๐˜€๐˜๐—ฒ๐—บ ๐—œ๐—บ๐—ฝ๐—น๐—ฒ๐—บ๐—ฒ๐—ป๐˜๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ฎ๐˜ ๐—–๐—ง๐—ก๐—Ÿ๐—ฃ๐—ฅ

Current capabilities include:

โœ… Named Entity Recognition
โœ… Entity Normalization
โœ… Pronoun Resolution
โœ… Possessive Resolution
โœ… Location Resolution
โœ… Rule-Based Noun Phrase Resolution
โœ… Coreference Chain Construction
โœ… Streamlit-Based Visualization

The system acts as a foundational layer between entity extraction and knowledge graph generation.

๐—™๐˜‚๐˜๐˜‚๐—ฟ๐—ฒ ๐—ช๐—ผ๐—ฟ๐—ธ

โ€ข Multi-entity discourse memory
โ€ข Entity salience tracking
โ€ข Advanced noun phrase resolution
โ€ข Relation-aware coreference resolution
โ€ข Knowledge graph integration
โ€ข Hybrid neural-rule architectures

๐—–๐—ผ๐—ป๐—ฐ๐—น๐˜‚๐˜€๐—ถ๐—ผ๐—ป

Building effective Tamil Information Extraction systems requires more than Named Entity Recognition.

By introducing a dedicated coreference resolution layer, we can maintain entity consistency across documents, improve relation extraction quality, and generate more reliable structured knowledge.

For low-resource languages such as Tamil, carefully designed rule-based systems remain a practical and effective pathway toward document-level semantic understanding while larger neural approaches continue to mature.

โšก๏ธ๐—™๐—ถ๐—ป๐—ฒ-๐—ง๐˜‚๐—ป๐—ถ๐—ป๐—ด ๐—œ๐—ป๐—ฑ๐—ถ๐—ฐ๐—ก๐—˜๐—ฅ ๐—ณ๐—ผ๐—ฟ ๐—ฆ๐—ฟ๐—ถ ๐—Ÿ๐—ฎ๐—ป๐—ธ๐—ฎ๐—ป ๐—ง๐—ฎ๐—บ๐—ถ๐—น ๐—ก๐—ฎ๐—บ๐—ฒ๐—ฑ ๐—˜๐—ป๐˜๐—ถ๐˜๐˜† ๐—ฅ๐—ฒ๐—ฐ๐—ผ๐—ด๐—ป๐—ถ๐˜๐—ถ๐—ผ๐—ปTransformer-based multilingual NLP systems have sign...
03/06/2026

โšก๏ธ๐—™๐—ถ๐—ป๐—ฒ-๐—ง๐˜‚๐—ป๐—ถ๐—ป๐—ด ๐—œ๐—ป๐—ฑ๐—ถ๐—ฐ๐—ก๐—˜๐—ฅ ๐—ณ๐—ผ๐—ฟ ๐—ฆ๐—ฟ๐—ถ ๐—Ÿ๐—ฎ๐—ป๐—ธ๐—ฎ๐—ป ๐—ง๐—ฎ๐—บ๐—ถ๐—น ๐—ก๐—ฎ๐—บ๐—ฒ๐—ฑ ๐—˜๐—ป๐˜๐—ถ๐˜๐˜† ๐—ฅ๐—ฒ๐—ฐ๐—ผ๐—ด๐—ป๐—ถ๐˜๐—ถ๐—ผ๐—ป

Transformer-based multilingual NLP systems have significantly improved Named Entity Recognition (NER) across many languages. However, low-resource language variants such as Sri Lankan Tamil still face substantial challenges due to limited domain-specific datasets and linguistic underrepresentation.

At CTNLPR, we fine-tuned ๐—ฎ๐—ถ๐Ÿฐ๐—ฏ๐—ต๐—ฎ๐—ฟ๐—ฎ๐˜/๐—œ๐—ป๐—ฑ๐—ถ๐—ฐ๐—ก๐—˜๐—ฅ specifically for Sri Lankan Tamil using a custom annotated NER corpus.

๐—ข๐—ฏ๐—ท๐—ฒ๐—ฐ๐˜๐—ถ๐˜ƒ๐—ฒ

Improve entity recognition for:

โ€ข Sri Lankan Tamil linguistic patterns
โ€ข Local person, location, and organization names
โ€ข Morphology-aware contextual variations

๐—ช๐—ต๐˜† ๐—ฆ๐—ฟ๐—ถ ๐—Ÿ๐—ฎ๐—ป๐—ธ๐—ฎ๐—ป ๐—ง๐—ฎ๐—บ๐—ถ๐—น ๐—ก๐—˜๐—ฅ ๐—ถ๐˜€ ๐—–๐—ต๐—ฎ๐—น๐—น๐—ฒ๐—ป๐—ด๐—ถ๐—ป๐—ด

Most multilingual NER systems are trained primarily on:

โ€ข General web corpora
โ€ข Indian Tamil datasets
โ€ข Multilingual benchmark datasets
โ€ข Formal textual sources

When applied to Sri Lankan Tamil, they often struggle with:

โ€ข Regional naming conventions
โ€ข Local organization terminology
โ€ข Morphological suffix complexity
โ€ข OCR-induced token inconsistencies
โ€ข Subword tokenization fragmentation
โ€ข Ambiguous entity boundaries

These limitations directly affect downstream systems such as:

โ€ข Semantic Search
โ€ข Document Intelligence
โ€ข Knowledge Graph Construction
โ€ข Tamil Chatbots
โ€ข RAG Systems
โ€ข Government Document Processing

๐— ๐—ผ๐—ฑ๐—ฒ๐—น ๐—™๐—ถ๐—ป๐—ฒ-๐—ง๐˜‚๐—ป๐—ถ๐—ป๐—ด ๐—ข๐˜ƒ๐—ฒ๐—ฟ๐˜ƒ๐—ถ๐—ฒ๐˜„

๐—•๐—ฎ๐˜€๐—ฒ ๐— ๐—ผ๐—ฑ๐—ฒ๐—น

โ†’ ai4bharat/IndicNER

๐—˜๐—ป๐˜๐—ถ๐˜๐˜† ๐—ง๐˜†๐—ฝ๐—ฒ๐˜€

โ€ข PERSON
โ€ข LOCATION
โ€ข ORGANIZATION

๐—ž๐—ฒ๐˜† ๐—ข๐—ฝ๐˜๐—ถ๐—บ๐—ถ๐˜‡๐—ฎ๐˜๐—ถ๐—ผ๐—ป๐˜€

โœ… Tamil-safe Tokenization
โœ… Unicode Normalization
โœ… BIO Tagging
โœ… Proper Subword Label Alignment
โœ… Morphology-aware Training
โœ… OCR-aware Preprocessing

๐—ง๐—ฒ๐—ฐ๐—ต๐—ป๐—ถ๐—ฐ๐—ฎ๐—น ๐—–๐—ต๐—ฎ๐—น๐—น๐—ฒ๐—ป๐—ด๐—ฒ๐˜€

1๏ธโƒฃ ๐—ง๐—ฎ๐—บ๐—ถ๐—น ๐—ง๐—ผ๐—ธ๐—ฒ๐—ป๐—ถ๐˜‡๐—ฎ๐˜๐—ถ๐—ผ๐—ป

Tamil is morphologically rich. Incorrect tokenization can cause:

โ€ข Broken entity spans
โ€ข Incorrect BIO labels
โ€ข Fragmented predictions

2๏ธโƒฃ ๐—ฆ๐˜‚๐—ฏ๐˜„๐—ผ๐—ฟ๐—ฑ ๐—Ÿ๐—ฎ๐—ฏ๐—ฒ๐—น ๐—”๐—น๐—ถ๐—ด๐—ป๐—บ๐—ฒ๐—ป๐˜

Transformer tokenizers frequently split Tamil words into multiple subword units.

Without proper alignment:

โ€ข Entity spans become corrupted
โ€ข BIO labels mismatch
โ€ข Training instability increases

3๏ธโƒฃ ๐—ข๐—–๐—ฅ ๐—ก๐—ผ๐—ถ๐˜€๐—ฒ

Tamil OCR systems still generate:

โ€ข Grapheme inconsistencies
โ€ข Merged tokens
โ€ข Invalid Unicode combinations
โ€ข Punctuation corruption

Therefore OCR-aware normalization was integrated before training.

๐— ๐—ผ๐—ฑ๐—ฒ๐—น ๐—˜๐˜ƒ๐—ฎ๐—น๐˜‚๐—ฎ๐˜๐—ถ๐—ผ๐—ป

๐—ข๐˜ƒ๐—ฒ๐—ฟ๐—ฎ๐—น๐—น ๐—ฃ๐—ฒ๐—ฟ๐—ณ๐—ผ๐—ฟ๐—บ๐—ฎ๐—ป๐—ฐ๐—ฒ

โ€ข F1 Score โ†’ 0.650
โ€ข Precision โ†’ 0.602
โ€ข Recall โ†’ 0.707
โ€ข Accuracy โ†’ 96.04%

๐—˜๐—ป๐˜๐—ถ๐˜๐˜†-๐˜„๐—ถ๐˜€๐—ฒ ๐—™๐Ÿญ

โ€ข PERSON โ†’ 0.721
โ€ข LOCATION โ†’ 0.698
โ€ข ORGANIZATION โ†’ 0.484

PERSON and LOCATION categories achieved relatively strong performance, while ORGANIZATION entities remain the most challenging category.

๐—œ๐—ป๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—˜๐˜…๐—ฎ๐—บ๐—ฝ๐—น๐—ฒ๐˜€

๐—˜๐˜…๐—ฎ๐—บ๐—ฝ๐—น๐—ฒ ๐Ÿญ

Sentence:
"เฎชเฎพเฎฐเฎคเฎฟเฎคเฎพเฎšเฎฉเฏ เฎŽเฎดเฏเฎคเฎฟเฎฏ เฎจเฏ‚เฎฒเฏˆ เฎชเฎพเฎฐเฎคเฎฟ เฎชเฎคเฎฟเฎชเฏเฎชเฎ•เฎฎเฏ เฎตเฏ†เฎณเฎฟเฎฏเฎฟเฎŸเฏเฎŸเฎคเฏ."

Output:
๐Ÿ‘ค PERSON โ†’ เฎชเฎพเฎฐเฎคเฎฟเฎคเฎพเฎšเฎฉเฏ
๐Ÿข ORGANIZATION โ†’ เฎชเฎพเฎฐเฎคเฎฟ เฎชเฎคเฎฟเฎชเฏเฎชเฎ•เฎฎเฏ

๐—˜๐˜…๐—ฎ๐—บ๐—ฝ๐—น๐—ฒ ๐Ÿฎ

Sentence:
"เฎตเฎŸเฎฎเฎฐเฎพเฎŸเฏเฎšเฎฟ เฎคเฏŠเฎดเฎฟเฎฒเฏเฎจเฏเฎŸเฏเฎช เฎจเฎฟเฎฑเฏเฎตเฎฉเฎฎเฏ เฎฎเฎพเฎฃเฎตเฎฐเฏเฎ•เฎณเฏˆ เฎšเฏ‡เฎฐเฏเฎคเฏเฎคเฎคเฏ."

Output:
๐Ÿข ORGANIZATION โ†’ เฎตเฎŸเฎฎเฎฐเฎพเฎŸเฏเฎšเฎฟ เฎคเฏŠเฎดเฎฟเฎฒเฏเฎจเฏเฎŸเฏเฎช เฎจเฎฟเฎฑเฏเฎตเฎฉเฎฎเฏ

๐—˜๐˜…๐—ฎ๐—บ๐—ฝ๐—น๐—ฒ ๐Ÿฏ

Sentence:
"เฎจเฎตเฎฎเฎฃเฎฟ เฎ•เฎฟเฎฐเฎพเฎฎเฎฎเฏ เฎตเฏ†เฎณเฏเฎณเฎคเฏเฎคเฎพเฎฒเฏ เฎชเฎพเฎคเฎฟเฎ•เฏเฎ•เฎชเฏเฎชเฎŸเฏเฎŸเฎคเฏ."

Output:
๐Ÿ“ LOCATION โ†’ เฎจเฎตเฎฎเฎฃเฎฟ

๐—˜๐˜…๐—ฎ๐—บ๐—ฝ๐—น๐—ฒ ๐Ÿฐ

Sentence:
"เฎœเฏ‡.เฎ.เฎŽเฎธเฏ.เฎชเฏ€. เฎœเฎฏเฎšเฎฟเฎ™เฏเฎ• เฎจเฎตเฎฎเฎฃเฎฟ เฎ•เฎฟเฎฐเฎพเฎฎเฎคเฏเฎคเฎฟเฎฑเฏเฎ•เฏ เฎšเฏ†เฎฉเฏเฎฑเฎพเฎฐเฏ."

Output:
๐Ÿ‘ค PERSON โ†’ เฎœเฏ‡.เฎ.เฎŽเฎธเฏ.เฎชเฏ€. เฎœเฎฏเฎšเฎฟเฎ™เฏเฎ•
๐Ÿ“ LOCATION โ†’ เฎจเฎตเฎฎเฎฃเฎฟ

๐—ž๐—ฒ๐˜† ๐—ข๐—ฏ๐˜€๐—ฒ๐—ฟ๐˜ƒ๐—ฎ๐˜๐—ถ๐—ผ๐—ป

One of the most important findings from this work is:

"Better preprocessing and domain-specific data can be as important as model architecture."

For low-resource languages like Sri Lankan Tamil:

โ€ข High-quality annotations matter
โ€ข OCR normalization matters
โ€ข Tokenizer alignment matters
โ€ข Linguistic preprocessing matters

Large transformer architectures alone are not sufficient without carefully prepared language-specific datasets.

๐—”๐—ฝ๐—ฝ๐—น๐—ถ๐—ฐ๐—ฎ๐˜๐—ถ๐—ผ๐—ป๐˜€

โ€ข Tamil NER Systems
โ€ข Semantic Search
โ€ข RAG Pipelines
โ€ข OCR Information Extraction
โ€ข Knowledge Graph Construction
โ€ข Tamil Chatbots

This work is part of ongoing Tamil NLP research at CTNLPR aimed at building stronger NLP infrastructure for low-resource Tamil language technologies.

๐ŸŒŸ๐๐ฎ๐ข๐ฅ๐๐ข๐ง๐  ๐š ๐’๐ซ๐ข ๐‹๐š๐ง๐ค๐š๐ง ๐“๐š๐ฆ๐ข๐ฅ ๐๐š๐ฆ๐ž๐ ๐„๐ง๐ญ๐ข๐ญ๐ฒ ๐‘๐ž๐œ๐จ๐ ๐ง๐ข๐ญ๐ข๐จ๐ง ๐ƒ๐š๐ญ๐š๐ฌ๐ž๐ญ ๐Ÿ๐จ๐ซ ๐‹๐จ๐ฐ-๐‘๐ž๐ฌ๐จ๐ฎ๐ซ๐œ๐ž ๐๐‹๐The growth of Large Language Models (L...
25/05/2026

๐ŸŒŸ๐๐ฎ๐ข๐ฅ๐๐ข๐ง๐  ๐š ๐’๐ซ๐ข ๐‹๐š๐ง๐ค๐š๐ง ๐“๐š๐ฆ๐ข๐ฅ ๐๐š๐ฆ๐ž๐ ๐„๐ง๐ญ๐ข๐ญ๐ฒ ๐‘๐ž๐œ๐จ๐ ๐ง๐ข๐ญ๐ข๐จ๐ง ๐ƒ๐š๐ญ๐š๐ฌ๐ž๐ญ ๐Ÿ๐จ๐ซ ๐‹๐จ๐ฐ-๐‘๐ž๐ฌ๐จ๐ฎ๐ซ๐œ๐ž ๐๐‹๐

The growth of Large Language Models (LLMs) and multilingual NLP systems has significantly improved language technologies across major global languages. However, low-resource languages such as Sri Lankan Tamil still face a severe lack of high-quality annotated datasetsโ€”especially for foundational tasks like Named Entity Recognition (NER).

To address this gap, we developed the Srilankan-Tamil-NER Dataset, a Tamil NER dataset designed specifically for Sri Lankan Tamil linguistic and contextual usage.

This dataset is intended to support:

โ€ข Tamil NER research
โ€ข Indic language fine-tuning
โ€ข Information extraction systems
โ€ข Retrieval-Augmented Generation (RAG)
โ€ข Tamil LLM adaptation
โ€ข Domain-specific AI systems for Sri Lanka

๐—ช๐—ต๐˜† ๐—ฆ๐—ฟ๐—ถ ๐—Ÿ๐—ฎ๐—ป๐—ธ๐—ฎ๐—ป ๐—ง๐—ฎ๐—บ๐—ถ๐—น ๐—ก๐—˜๐—ฅ ๐— ๐—ฎ๐˜๐˜๐—ฒ๐—ฟ๐˜€

Named Entity Recognition (NER) is a core NLP task that identifies and classifies entities such as:

โ€ข Person names
โ€ข Locations
โ€ข Organizations
โ€ข Dates
โ€ข Miscellaneous entities

NER acts as a foundational layer for many downstream NLP systems including:

โ€ข Question answering
โ€ข Search systems
โ€ข Chatbots
โ€ข Document intelligence
โ€ข Machine translation
โ€ข Knowledge graph generation

For Tamil โ€” particularly Sri Lankan Tamil โ€” publicly available annotated corpora remain extremely limited. Existing multilingual datasets often underrepresent regional linguistic variations, local named entities, and culturally contextual terminology.

Most existing NER systems for Tamil are trained on datasets originating from Indian Tamil corpora, leaving significant gaps in handling:

โ€ข Sri Lankan Tamil vocabulary
โ€ข Local organization names
โ€ข Sri Lankan place names
โ€ข Government and institutional terminology

Our dataset aims to bridge this gap.

๐—”๐—ฏ๐—ผ๐˜‚๐˜ ๐˜๐—ต๐—ฒ ๐——๐—ฎ๐˜๐—ฎ๐˜€๐—ฒ๐˜

๐——๐—ฎ๐˜๐—ฎ๐˜€๐—ฒ๐˜ ๐—ก๐—ฎ๐—บ๐—ฒ:
Srilankan-Tamil-NER Dataset

The primary goal of this dataset is to create a high-quality manually curated Named Entity Recognition corpus for Sri Lankan Tamil under CTNLPR.

The dataset is structured to support fine-tuning transformer-based multilingual models such as:

โ€ข IndicNER
โ€ข mBERT
โ€ข XLM-RoBERTa
โ€ข MuRIL
โ€ข IndicBERT

๐——๐—ฎ๐˜๐—ฎ๐˜€๐—ฒ๐˜ ๐—ฆ๐˜๐—ฎ๐˜๐—ถ๐˜€๐˜๐—ถ๐—ฐ๐˜€

โ€ข B-PER (Person): 4,533
โ€ข B-LOC (Location): 8,110
โ€ข B-ORG (Organization): 3,369
โ€ข Total Entities: 16,012

๐——๐—ฎ๐˜๐—ฎ๐˜€๐—ฒ๐˜ ๐—ฃ๐—ฟ๐—ฒ๐—ฝ๐—ฎ๐—ฟ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ฃ๐—ถ๐—ฝ๐—ฒ๐—น๐—ถ๐—ป๐—ฒ

Creating a Tamil NER dataset involves significantly more than simple annotation.

The preparation workflow included multiple stages:

1. ๐‘ซ๐’‚๐’•๐’‚ ๐‘ช๐’๐’๐’๐’†๐’„๐’•๐’Š๐’๐’

The raw Tamil text corpus was collected from the Noolaham corpus and other relevant publicly available Sri Lankan Tamil textual sources.

Special attention was given to:

โ€ข Local linguistic relevance
โ€ข Entity diversity
โ€ข Sentence quality
โ€ข Contextual richness

The objective was to capture realistic Sri Lankan Tamil usage patterns rather than synthetic or translated text.

2. ๐‘ถ๐‘ช๐‘น ๐’‚๐’๐’… ๐‘ป๐’†๐’™๐’• ๐‘ต๐’๐’“๐’Ž๐’‚๐’๐’Š๐’›๐’‚๐’•๐’Š๐’๐’

Tamil NLP pipelines often begin with scanned or image-based documents.

As part of our broader Tamil document intelligence workflow, OCR-extracted Tamil text underwent:

โ€ข Unicode normalization
โ€ข Punctuation cleaning
โ€ข Whitespace normalization
โ€ข Invalid character filtering
โ€ข OCR noise reduction

OCR-related preprocessing becomes extremely important because Tamil script errors can propagate heavily into token classification systems.

3. ๐‘ต๐’‚๐’Ž๐’†๐’… ๐‘ฌ๐’๐’•๐’Š๐’•๐’š ๐‘จ๐’๐’๐’๐’•๐’‚๐’•๐’Š๐’๐’

The dataset was manually annotated using BIO tagging format.

Entity Types:

โ€ข B-PER โ€” Beginning of person entity
โ€ข I-PER โ€” Inside person entity
โ€ข B-LOC โ€” Beginning of location entity
โ€ข I-LOC โ€” Inside location entity
โ€ข B-ORG โ€” Beginning of organization entity
โ€ข I-ORG โ€” Inside organization entity
โ€ข O โ€” Non-entity token

Example:

เฎ‡เฎฐเฎพเฎฎเฎจเฎพเฎคเฎฉเฏ โ†’ B-PER
เฎฏเฎพเฎดเฏเฎชเฏเฎชเฎพเฎฃเฎฎเฏ โ†’ B-LOC
เฎชเฎฒเฏเฎ•เฎฒเฏˆเฎ•เฏเฎ•เฎดเฎ•เฎฎเฏ โ†’ B-ORG

๐—–๐—ต๐—ฎ๐—น๐—น๐—ฒ๐—ป๐—ด๐—ฒ๐˜€ ๐—ถ๐—ป ๐—ฆ๐—ฟ๐—ถ ๐—Ÿ๐—ฎ๐—ป๐—ธ๐—ฎ๐—ป ๐—ง๐—ฎ๐—บ๐—ถ๐—น ๐—ก๐—˜๐—ฅ

Building a Tamil NER dataset introduced several language-specific challenges.

โ€ข Morphological complexity
โ€ข OCR noise
โ€ข Unicode inconsistencies
โ€ข Token boundary detection
โ€ข Subword alignment
โ€ข Limited benchmark corpora

๐—™๐—ถ๐—ป๐—ฒ-๐—ง๐˜‚๐—ป๐—ถ๐—ป๐—ด ๐—จ๐˜€๐—ฒ ๐—–๐—ฎ๐˜€๐—ฒ๐˜€

This dataset can support:

โ€ข Tamil NER
โ€ข OCR post-processing
โ€ข Semantic search systems
โ€ข RAG pipelines
โ€ข Tamil chatbots
โ€ข Government document AI
โ€ข Knowledge graph generation

The Srilankan-Tamil-NER Dataset, developed under CTNLPR, represents an important step toward strengthening the Sri Lankan Tamil NLP ecosystem through high-quality entity annotation and linguistically relevant corpus preparation.

#๐‘†๐‘Ÿ๐‘–๐ฟ๐‘Ž๐‘›๐‘˜๐‘Ž๐‘›๐‘‡๐‘Ž๐‘š๐‘–๐‘™ #๐‘‡๐‘Ž๐‘š๐‘–๐‘™๐‘๐ธ๐‘… #๐‘‡๐‘Ž๐‘š๐‘–๐‘™๐‘๐ฟ๐‘ƒ #๐‘๐‘Ž๐‘š๐‘’๐‘‘๐ธ๐‘›๐‘ก๐‘–๐‘ก๐‘ฆ๐‘…๐‘’๐‘๐‘œ๐‘”๐‘›๐‘–๐‘ก๐‘–๐‘œ๐‘› #๐ฟ๐‘œ๐‘ค๐‘…๐‘’๐‘ ๐‘œ๐‘ข๐‘Ÿ๐‘๐‘’๐‘๐ฟ๐‘ƒ #๐ผ๐‘›๐‘‘๐‘–๐‘๐‘๐ฟ๐‘ƒ #๐‘€๐‘ข๐‘…๐ผ๐ฟ #๐‘š๐ต๐ธ๐‘…๐‘‡ #๐‘‹๐ฟ๐‘€๐‘… #๐ต๐ผ๐‘‚๐‘ก๐‘Ž๐‘”๐‘”๐‘–๐‘›๐‘” #๐ฟ๐ฟ๐‘€ #๐‘…๐ด๐บ #๐‘†๐‘’๐‘š๐‘Ž๐‘›๐‘ก๐‘–๐‘๐‘†๐‘’๐‘Ž๐‘Ÿ๐‘โ„Ž #๐พ๐‘›๐‘œ๐‘ค๐‘™๐‘’๐‘‘๐‘”๐‘’๐บ๐‘Ÿ๐‘Ž๐‘โ„Ž๐‘  #๐ด๐ผ๐ธ๐‘›๐‘”๐‘–๐‘›๐‘’๐‘’๐‘Ÿ๐‘–๐‘›๐‘” #๐ท๐‘œ๐‘๐‘ข๐‘š๐‘’๐‘›๐‘ก๐ผ๐‘›๐‘ก๐‘’๐‘™๐‘™๐‘–๐‘”๐‘’๐‘›๐‘๐‘’ #๐ธ๐‘›๐‘ก๐‘–๐‘ก๐‘ฆ๐ธ๐‘ฅ๐‘ก๐‘Ÿ๐‘Ž๐‘๐‘ก๐‘–๐‘œ๐‘› #๐‘‚๐ถ๐‘… #๐‘‡๐‘Ÿ๐‘Ž๐‘›๐‘ ๐‘“๐‘œ๐‘Ÿ๐‘š๐‘’๐‘Ÿ๐‘€๐‘œ๐‘‘๐‘’๐‘™๐‘  #๐ถ๐‘œ๐‘š๐‘๐‘ข๐‘ก๐‘Ž๐‘ก๐‘–๐‘œ๐‘›๐‘Ž๐‘™๐ฟ๐‘–๐‘›๐‘”๐‘ข๐‘–๐‘ ๐‘ก๐‘–๐‘๐‘  #๐ถ๐‘‡๐‘๐ฟ๐‘ƒ๐‘…

โšก๏ธ ๐๐ฎ๐ข๐ฅ๐๐ข๐ง๐  ๐š ๐“๐š๐ฆ๐ข๐ฅ ๐‚๐จ๐ซ๐ž๐Ÿ๐ž๐ซ๐ž๐ง๐œ๐ž ๐‘๐ž๐ฌ๐จ๐ฅ๐ฎ๐ญ๐ข๐จ๐ง ๐’๐ฒ๐ฌ๐ญ๐ž๐ฆ: ๐™๐ž๐ซ๐จ-๐’๐ก๐จ๐ญ ๐‚๐จ๐ง๐ญ๐ž๐ฑ๐ญ๐ฎ๐š๐ฅ ๐Œ๐จ๐๐ž๐ฅ๐ข๐ง๐  ๐ฐ๐ข๐ญ๐ก ๐Œ๐ฎ๐‘๐ˆ๐‹Coreference Resolution (CR) i...
15/05/2026

โšก๏ธ ๐๐ฎ๐ข๐ฅ๐๐ข๐ง๐  ๐š ๐“๐š๐ฆ๐ข๐ฅ ๐‚๐จ๐ซ๐ž๐Ÿ๐ž๐ซ๐ž๐ง๐œ๐ž ๐‘๐ž๐ฌ๐จ๐ฅ๐ฎ๐ญ๐ข๐จ๐ง ๐’๐ฒ๐ฌ๐ญ๐ž๐ฆ: ๐™๐ž๐ซ๐จ-๐’๐ก๐จ๐ญ ๐‚๐จ๐ง๐ญ๐ž๐ฑ๐ญ๐ฎ๐š๐ฅ ๐Œ๐จ๐๐ž๐ฅ๐ข๐ง๐  ๐ฐ๐ข๐ญ๐ก ๐Œ๐ฎ๐‘๐ˆ๐‹

Coreference Resolution (CR) is a critical NLP task for identifying whether multiple mentions in a document refer to the same real-world entity. It plays a major role in:

โ€ข Knowledge Graph Construction
โ€ข Relation Extraction
โ€ข Semantic Search
โ€ข RAG Systems
โ€ข Conversational AI
โ€ข Document-Level Understanding

While English NLP already has mature coreference systems and libraries, Tamil remains a highly challenging low-resource language for discourse-level semantic modeling.

At CTNLPR, we explored how modern multilingual coreference architectures can be adapted for Tamil using zero-shot contextual semantic modeling instead of heavily supervised pipelines.

Tamil introduces several difficult linguistic challenges:

โ€ข Agglutinative morphology
โ€ข Free word order
โ€ข Pronoun dropping
โ€ข Implicit subject references
โ€ข Rich inflectional structures
โ€ข Long-distance discourse dependencies
โ€ข Noun-to-noun semantic references

Additionally:

โ€ข No dedicated Tamil coreference libraries currently exist publicly
โ€ข Large annotated Tamil CR datasets are unavailable
โ€ข Most multilingual systems remain heavily English-biased

โš™๏ธ Our Architecture

The architecture currently being explored at CTNLPR uses:

โ€ข MuRIL-based contextual embeddings
โ€ข Span-based mention detection
โ€ข Contextual span representations
โ€ข Cosine similarity-based semantic linking
โ€ข Agglomerative clustering

Instead of manually defining antecedents, the system automatically generates semantic mention spans from Tamil text and groups semantically related mentions into discourse-level entity chains using contextual similarity.

๐Ÿง  Key Technical Direction

Traditional supervised coreference systems depend heavily on:

โ€ข Large annotated corpora
โ€ข Expensive training pipelines
โ€ข Language-specific supervision
โ€ข Antecedent ranking architectures
โ€ข High computational cost

For Tamil, these resources are extremely limited.

Our approach avoids heavy annotation dependency while still leveraging multilingual transformer-based semantic understanding learned from Indian-language pretraining.

Each candidate span is encoded using contextual embeddings generated from MuRIL, and span-level semantic representations are constructed using contextual token pooling. The system then performs:

โ€ข Heuristic span pruning
โ€ข Semantic similarity computation
โ€ข Similarity-driven clustering

to generate discourse-level coreference chains.

๐Ÿš€ Key Advantages

โ€ข Zero-shot inference
โ€ข Low-resource scalability
โ€ข Context-aware semantic reasoning
โ€ข Better adaptation to Tamil morphology
โ€ข Lightweight unsupervised inference
โ€ข Reduced annotation dependency

This architecture is being explored at CTNLPR as a foundation for:

โ€ข Tamil discourse understanding
โ€ข Entity-aware semantic linking
โ€ข Knowledge Graph Construction
โ€ข Ontology-aware NLP
โ€ข Multilingual semantic reasoning
โ€ข Advanced RAG systems

๐Ÿ”ฌ Building document-level semantic understanding for Tamil is one of the next major steps toward scalable low-resource AI systems.

โšก๐๐ฎ๐ข๐ฅ๐๐ข๐ง๐  ๐š ๐Œ๐จ๐ซ๐ฉ๐ก๐จ๐ฅ๐จ๐ ๐ฒ-๐€๐ฐ๐š๐ซ๐ž ๐“๐š๐ฆ๐ข๐ฅ ๐๐„๐‘ ๐๐ข๐ฉ๐ž๐ฅ๐ข๐ง๐ž: ๐…๐ซ๐จ๐ฆ ๐“๐ซ๐š๐ง๐ฌ๐Ÿ๐จ๐ซ๐ฆ๐ž๐ซ ๐„๐ฑ๐ญ๐ซ๐š๐œ๐ญ๐ข๐จ๐ง ๐ญ๐จ ๐‚๐š๐ง๐จ๐ง๐ข๐œ๐š๐ฅ ๐„๐ง๐ญ๐ข๐ญ๐ฒ ๐‘๐ž๐ฌ๐จ๐ฅ๐ฎ๐ญ๐ข๐จ๐งNamed Entity ...
04/05/2026

โšก๐๐ฎ๐ข๐ฅ๐๐ข๐ง๐  ๐š ๐Œ๐จ๐ซ๐ฉ๐ก๐จ๐ฅ๐จ๐ ๐ฒ-๐€๐ฐ๐š๐ซ๐ž ๐“๐š๐ฆ๐ข๐ฅ ๐๐„๐‘ ๐๐ข๐ฉ๐ž๐ฅ๐ข๐ง๐ž: ๐…๐ซ๐จ๐ฆ ๐“๐ซ๐š๐ง๐ฌ๐Ÿ๐จ๐ซ๐ฆ๐ž๐ซ ๐„๐ฑ๐ญ๐ซ๐š๐œ๐ญ๐ข๐จ๐ง ๐ญ๐จ ๐‚๐š๐ง๐จ๐ง๐ข๐œ๐š๐ฅ ๐„๐ง๐ญ๐ข๐ญ๐ฒ ๐‘๐ž๐ฌ๐จ๐ฅ๐ฎ๐ญ๐ข๐จ๐ง

Named Entity Recognition (NER) is a critical layer in our Tamil NLP stack (search, indexing, knowledge graph construction, and RAG).
However, for Tamil, extracting entities is only half the problem โ€” canonicalizing them is the real challenge.

๐Ÿ’ฅ What We Built in CTNLPR

We designed a Tamil-aware NER pipeline by extending a transformer-based model with morphological normalization:

โ€ข ai4bharat/IndicNER โ†’ baseline entity extraction
โ€ข Custom span merging โ†’ IOB consolidation
โ€ข Prefix-based grouping โ†’ variant clustering
โ€ข Morphological normalization layer โ†’ canonical entity resolution

Since Indic NER models are not morphology-aware, we explicitly evaluated and integrated normalization strategies.

๐Ÿ”ฐMethod Exploration (What We Tried)

โ€ข IndicNER (Transformer baseline)
โœ… Strong recall across entity types
โŽ Produces multiple inflected variants of the same entity

โ€ข Prefix-based grouping
โœ… Fast heuristic clustering
โŽ Not linguistically grounded

โ€ข UoM Thamizhi Morphological Normalizer (University of Moratuwa)
โœ… Linguistically motivated rule-based approach
โŽ Limited effectiveness on real-world data
โŽ Struggled with:

* Noisy OCR text
* Complex suffix chains
* Unseen word forms

โ€ข Tamil Lemmatizer (final approach)
โœ… Consistent root-form extraction
โœ… Robust across inflected variants
โœ… Best empirical performance in our pipeline

๐Ÿ”ฌ Key Design Decision

Transformer models do not enforce canonical forms.

๐Ÿ‘‰ Surface forms like:
โ€ข เฎ‡เฎฒเฎ™เฏเฎ•เฏˆเฎฏเฎฟเฎฒเฏ
โ€ข เฎ‡เฎฒเฎ™เฏเฎ•เฏˆเฎฏเฎฟเฎฒเฏเฎฎเฏ
โ€ข เฎ‡เฎฒเฎ™เฏเฎ•เฏˆเฎฏเฎฟเฎฒเฏ‡

are extracted as separate entities

๐Ÿ‘‰ After normalization:
โ€ข เฎ‡เฎฒเฎ™เฏเฎ•เฏˆ

This enables many-to-one mapping, critical for system consistency.

๐Ÿงฉ Our Setup

Pipeline:

โ€ข Document โ†’ chunking
โ€ข Transformer inference (IndicNER)
โ€ข IOB span merging + filtering
โ€ข Variant aggregation (prefix-based)

โ€ข Morphological normalization (UoM explored โ†’ Lemmatizer selected)
โ€ข Entity re-indexing

Example:

เฎ‡เฎฒเฎ™เฏเฎ•เฏˆเฎฏเฎฟเฎฒเฏ โ†’ เฎ‡เฎฒเฎ™เฏเฎ•เฏˆ
เฎ‡เฎฒเฎ™เฏเฎ•เฏˆเฎฏเฎฟเฎฒเฏเฎฎเฏ โ†’ เฎ‡เฎฒเฎ™เฏเฎ•เฏˆ
เฎ‡เฎจเฏเฎคเฎฟเฎฏเฎพเฎตเฎฟเฎฒเฏ โ†’ เฎ‡เฎจเฏเฎคเฎฟเฎฏเฎพ

โžก๏ธ System-Level Challenges We Solved

โ€ข Agglutinative suffix handling
โ€ข Variant explosion in entity outputs
โ€ข OCR/noisy input robustness
โ€ข Canonical entity consistency across documents

๐Ÿ“Š What We Observed

โ€ข Transformer NER โ†’ high recall, low canonical consistency
โ€ข UoM morphological normalizer โ†’ linguistically sound but limited robustness
โ€ข Lemmatizer โ†’ best normalization performance in practice

๐ŸŒŸ Final system:
IndicNER + Lemmatization (hybrid architecture)

โœณ๏ธ Key Insight

In Tamil NER, the challenge is not detection โ€”
it is morphological normalization.

NER output โ‰  final entity

๐ŸŒŸ Canonicalization is essential for:
โ€ข Indexing
โ€ข Entity linking
โ€ข Knowledge graphs
โ€ข RAG systems

๐Ÿš€ Outcome

We built a production-ready Tamil NER system that:

โ€ข Resolves inflected entity variants
โ€ข Produces stable canonical forms
โ€ข Improves downstream retrieval and analytics
โ€ข Scales across multi-document pipelines

๐Ÿ”ฌ This work is part of ongoing Tamil NLP system development at CTNLPR

Address

63, Sir Pon, Thirunelvelly, Ramanathan Road, Kallady
Jaffna
40000

Opening Hours

Monday 09:00 - 17:00
Tuesday 09:00 - 17:00
Wednesday 09:00 - 17:00
Thursday 09:00 - 17:00
Friday 09:00 - 17:00

Telephone

+442037733854

Alerts

Be the first to know and let us send you an email when Center for Tamil Natural Language Processing Research posts news and promotions. Your email address will not be used for any other purpose, and you can unsubscribe at any time.

Contact The School

Send a message to Center for Tamil Natural Language Processing Research:

Shortcuts

Share

Category