OIHIK Lab · AI research & engineering

Research.
Build. Evaluate.
For everyone.

Advancing intelligent systems through better data, careful evaluation, and reproducible engineering. Connecting AI research with real needs in education, industry, and public service.

Machine learningLanguage AIOpen research

The research ecosystemFIG. 01
A connected research workflow

Reliable research starts with well-sourced corpora, documented data, and thoughtful preparation.

Workflow illustration

Research with relevance.
A shared foundation for progress.

Big questions.
Rigorous foundations.

From raw data to evaluated models, our research spans the full machine learning lifecycle. Each discipline informs the next.

The research lifecycle

Every result is a new starting point ↺
  1. 01

    Collect

    Corpora, documents, and source records

  2. 02

    Clean

    Normalization and deduplication

  3. 03

    Filter

    Quality checks and relevance

  4. 04

    Prepare

    Tokenization and metadata

  5. 05

    Train

    Pretraining and fine-tuning

  6. 06

    Evaluate

    Benchmarks and task evaluation

  7. 07

    Analyze

    Errors, metrics, and comparisons

  8. 08

    Iterate

    Refine the question and repeat

Better data.
Better questions. Better AI.

Our research interests connect model development with the data behind it, with a focus on Bangla, multilingual understanding, and useful evaluation.

Models

Architectures, adaptation, and experimentation.

  • Adapted models

    Fine-tuning for language, tasks, and domain needs.

    SFT
  • Experimental architectures

    Representation learning and controlled comparisons.

    RESEARCH

Datasets

The foundation of every meaningful experiment.

  • Bangla language resources

    Text, grammar, examples, and language knowledge.

    BANGLA
  • Multilingual corpora

    Bangla, English, and cross-language research data.

    CORPUS
  • Domain & evaluation data

    Curated sources, filtered data, and task benchmarks.

    QUALITY

Release documentation will include source information, versions, intended uses, and known limitations.

Measure. Compare.
Understand.

A model is more than a headline score. Our evaluation approach asks what works, where it fails, and whether another researcher can reproduce the result.

Perplexity · PPL
How well a model predicts text
Bits per byte · BPB
Prediction measured at the byte level
Task benchmarks
Performance on defined tasks
Token coverage
How language is represented
Data quality
Source quality and duplication
Error analysis
Failure patterns and limitations
Evaluation frameworkFIG. 02

Evidence before conclusions

Method overview / evaluation criteria

Research questionEvaluation lens
How well does it predict text?PPL / BPB
Can it perform the task?Task benchmarks
Is the data suitable?Quality / coverage
Where does it struggle?Error analysis
Can we compare fairly?Matched settings
  1. 01 / Configure
  2. 02 / Evaluate
  3. 03 / Review
Document the model, dataset, settings, and limits of every comparison.

Good research needs
great foundations.

A connected research environment brings data, compute, models, and evaluation into the same workflow. Trace the work from its source to its findings.

Reference architecture / source to insight
  1. 01 / InputData sources
  2. 02 / ProcessPreparation
  3. 03 / DataDataset versions
  4. 04 / ComputeTraining
  5. 05 / ModelsModel registry
  6. 06 / MeasureEvaluation
  7. 07 / AnalyzeResults
  8. 08 / LearnResearch
PythonGPU computeStorageAPIsExperiment trackingAutomation

From questions
to working foundations.

Language resources are a starting point. Data quality and reproducible evaluation shape the research directions around them.

01 / Language resource workspace

Language, carefully organized.

Collecting grammar rules, useful examples, and vocabulary in a structured workspace for language study.

LanguageStructured data
Open language workspace
02 / Research direction

Quality before quantity.

Exploring repeatable approaches to corpus preparation, filtering, deduplication, and source documentation.

Dataset engineeringQuality
Explore the data focus
03 / Research direction

Evidence you can examine.

Developing an evaluation approach that connects metrics with model behavior, context, and limitations.

BenchmarkingReproducibility
Explore the approach

Technical depth.
Human relevance.

AI research should be understandable beyond the lab. We connect technical questions with the needs of people, institutions, and the communities they serve.

Academia & researchers

Research questions, documented methods, and reproducible experiments for students, educators, universities, and fellow researchers.

Learning / Experimentation / Shared knowledge

Government & public sector

Research relevant to ministries, departments, local government, and public institutions: Bangla information access, document understanding, and evaluation of AI for public services.

Language access / Public information / Evaluation

Industry & builders

Practical research directions in domain data, model adaptation, and evaluation to help engineering teams examine what works in their context.

Applied research / Data quality / Engineering

People & communities

Language resources and accessible explanations that make room for Bangla speakers, independent learners, and open research communities.

Inclusion / Local language / Participation

Public value starts with careful research.Data provenance, privacy, language coverage, transparent evaluation, and human oversight are questions to examine throughout the research lifecycle.

Build. Measure.
Understand.
Improve.

Meaningful AI research takes more than training a model. It takes reliable data, repeatable experiments, careful evaluation, and the engineering to keep asking better questions.

  1. 01 / Trace the work

    Reproducibility

    Keep experiments repeatable, with clear configurations, data versions, and methods.

  2. 02 / Share the learning

    Open research

    Share knowledge, tools, and findings whenever rights and responsibilities allow.

  3. 03 / Start at the source

    Data quality

    Understand the origin, coverage, and limitations of the data behind every result.

  4. 04 / Build with care

    Engineering depth

    Treat infrastructure and reliable workflows as part of the research itself.

Different disciplines.
Shared curiosity.

A research lab is a meeting point for model researchers, data engineers, and people who understand language and real-world needs. These are the perspectives behind our approach.

Research perspective

Machine learning & evaluation

Asking precise questions about how intelligent systems learn and behave.

Engineering perspective

Data & research systems

Connecting reliable data preparation with traceable experimental workflows.

Domain perspective

Language & public needs

Bringing context, linguistic knowledge, and practical needs into the research.

Individual researcher profiles and publications are not yet listed.

Better AI starts
with better research.

Explore the ideas, data, and engineering behind OIHIK Lab. Find the questions that connect our research with your world.