Insights · Language
Ask the way you speak, with both languages in one sentence.
Super Intelligence (SI) that only works in textbook Bangla would fail on its first morning in a Dhaka office. This article is about the questions people really type, and how Bahlul SI handles them and is scored on them.
How people ask
Nobody types textbook Bangla
Listen to a branch for an hour. 'Circular 12 er late fee koto?' 'Leave policy te maternity leave koto din?' One sentence holds an English noun in a Bangla frame, a Bangla word in Latin letters and a Latin digit where a Bangla one might be. Help-desk logs look the same. So do the documents: an older circular in a legacy font, a scanned page with a stamp over the figure, lakh and crore beside millions, a Bangla calendar date next to a Gregorian one, and the same word spelled three ways.
A system built for clean text in one language meets this and guesses. Guessing is the one thing an assistant for circulars and SOPs must not do. The design goal is modest and strict: answer the question asked, in the user's language, from the documents the user may see, and say so when the documents do not hold the answer.
What the server does
Clean the question the way the documents were cleaned
When documents are loaded, text is extracted, scanned pages go through Bangla text recognition, legacy-font text is converted to Unicode, and the text is split into passages that keep the document name and the page. Passages are indexed two ways, by meaning and by keyword. A question gets the same treatment, so 'late fee' in Latin letters and the same words in Bangla script meet in the index. The system fetches the best passages this user is allowed to see, and the model answers from those passages only, citing each one, or says the documents do not answer.
Two rules keep the language straight. The answer follows the language of the question: a question in Bangla gets a Bangla answer, and a mixed question gets an answer in the mix the user wrote. Quoted passages stay in their original language, so a clause from an English policy is shown in English with the Bangla answer around it. Figures are kept as the document wrote them, and a figure is never converted from lakh to millions on the way.
The models
Open models, chosen by the test, switched by you
Bahlul SI runs open models, and the Bangla test bench chooses among them during the SI proof. A model that scores well in English and poorly on your scanned circulars loses to one that does the opposite. Many open models spend more tokens on Bangla than on the same text in English, which costs memory and speed, so the test bench measures speed in Bangla, not English. A public benchmark, BnMMLU, with 134,375 Bangla multiple-choice question-option pairs, found model results uneven and the gains from size diminishing (read 10 October 2026). That is why no model is trusted on its reputation.
You can switch models after a re-test. The licence of each model must allow your use, and it is recorded in the system card you receive at go-live. Where voice is wanted, open speech models for Bangla are tested the same way; if Bangla speech fails the test bench, we offer chat only and say so. Nothing is trained on your documents unless you agree to it in writing.
How it is scored
Eighty of 200 questions test the mess
In a 200-question test set, 25 questions are code-mixed or in Latin letters, 30 test numbers and dates, and 25 are answered only in scanned pages where text recognition broke letters or lines. Pass on a code-mixed item means the question asked was answered, in the user's language. Pass on a numbers item means every figure and date exactly right. Pass on a scanned item means the right answer, or a plain statement that the page cannot be read.
Two native Bangla reviewers score every answer on their own. Their agreement is measured, and the pass mark you approve, proposed at 85% on the critical questions with no critical failures, is fixed before anyone scores. The full method is in how the Bangla test bench scores an answer.
What stays with a person
A reader still decides
The system shows the page so a person can check it. An officer confirms any answer that changes what a customer or citizen is told. When the documents hold the answer three ways, the system shows the sources, but the choice of which rule applies today stays with the person whose job that is.
Spelling and script will drift as staff change. Under managed SI operations the full test set runs every month, which catches a model or a document set that has started to slip on code-mixed questions. Your own staff can add questions in their own words at any time; the real sentences from your help desk are the best test items there are.
FAQ
Questions we are asked
Will it answer in Bangla if I ask in English?
It answers in the language you asked in. Ask in English and you get English; ask in Bangla and you get Bangla; mix and it follows you. The quoted source stays as the document wrote it.
Does it read Bangla typed in Latin letters?
That is one of the six test categories, so it is measured on your documents rather than promised. If it scores poorly on your questions, the proof report says so before you buy a server.
Which model is it?
An open model chosen by the test bench on your documents during the SI proof, and named in the system card at go-live. You can switch it later after a re-test.
Send ten real questions
The best test of mixed-language handling is your own help-desk log. Bring ten real questions and ten documents that are not confidential to the first call, which is free.