Insights · Servers
One graphics card, one server room, no internet needed.
Super Intelligence (SI) on your own server means one rack server with one or two graphics cards inside it. This article says what that server holds, what it draws from the wall and what breaks, for the people who sign for it.
The box
What one server is
A Bahlul SI server is an ordinary rack server with a graphics card, a GPU, inside it. The card does the heavy arithmetic of the language model. The rest of the server holds the search index of your documents, the permissions, the log and the admin console. It sits in your server room or in a Bangladesh data centre, on your own network, with no route to the internet.
We size it in three configurations. Team is one 32 GB card for up to about 50 named users, about BDT 12 to 15 lakh of hardware. Department is one 48 GB data-centre card for a few hundred users, BDT 17 to 22 lakh. Company is two 96 GB cards for 1,000 or more users, BDT 55 to 65 lakh. These are indicative 2026 prices before duty and VAT. GPU prices moved sharply in 2026, so ask your supplier for a fresh quote. You buy the server directly; we specify it and check it. The servers page has the detail.
The limit
Memory on the card sets the limit
The card's memory decides two things: how big a model fits, and how many people can ask at the same moment. A model of 14 billion parameters at 8 bits takes about 17 GB. Each open conversation then takes a slice, about 0.8 GB for a 4,096-token context on a model that size. On a 48 GB card, after the model and the smaller helper models, that leaves room for about 27 people asking at once. If one named user in ten is active at any moment, that is about 270 named users.
That arithmetic comes from vendor sizing guides and from each model's own configuration file. It is an inference to test, not a promise. Memory shows whether the model fits, not whether answers come back fast enough. So no proposal quotes a user number until a load test has measured it, which is part of the three-week SI proof.
Three rules follow. Above the cap, requests queue and answers slow; the gateway's rate limits keep the service up. Doubling the context length roughly halves the users at once, which is why Bahlul SI sends the model only the best passages, never whole files. And many open models spend more tokens on Bangla than on the same text in English, so the test bench measures speed in Bangla, not English.
Power and heat
What it draws from the wall
Plan for about 1 kW for a one-card server at full load. The design assumption is a card of about 300 to 350 W, the rest of the server at about 600 W, and 100 W for the switch on the same supply. For scale, the maker's data sheet for one 48 GB data-centre card states a maximum of 350 W (NVIDIA L40S, read 10 October 2026). Your card may differ. Check the figures with your supplier's power calculator before the electrician sizes the circuit.
Every watt becomes heat. One watt is 3.412 BTU an hour, so a 1 kW server adds about 3,400 BTU an hour to the room, roughly a third of what a one-tonne air conditioner removes. A two-card Company server roughly doubles the card figure. Running cost follows the same number: at 730 hours a month, 1 kW is about 730 kWh. At a tariff of BDT 10 per kWh, an assumption to replace with your own bill, that is about BDT 7,300 a month before cooling. Mains here is 220 V, 50 Hz; confirm on site.
Power cuts
The UPS and the clean shutdown
Power cuts are a design condition, not an exception. The server sits behind an online double-conversion UPS with a network card. The UPS carries the load through a short cut. If power stays off, it tells the server to shut down cleanly after a set number of minutes, so the index and the databases are never left half-written. When power returns, the services restart in order on their own.
The sizing is simple. Load divided by a 0.9 power factor, times 1.25 headroom so the UPS never runs above 80%: a 1 kW load needs about 1,400 VA, so a 2 kVA unit for Team. Department is planned at 3 kVA, Company at 6 kVA with a generator. The design baseline is at least 15 minutes on battery at full load, sized for the end of the batteries' life and confirmed at the site survey. If you have a generator, the UPS only needs to bridge the start-up.
What breaks
One server, one point of failure
The design survives a power cut and a failed drive, because the drives are mirrored and backups run nightly to your own backup target. It has no standby server. A failed card or motherboard means downtime until the supplier repairs it under the three-year on-site warranty we ask you to buy, or a slower fallback on the processor. If your organisation cannot accept that, a second server is scoped separately.
Updates arrive monthly on hardware-encrypted USB drives with a manifest of checksums, installed in a maintenance window after a backup, with a rollback if any check fails. Nothing from your server goes out on them. Renting a comparable card abroad costs about USD 1.09 an hour, a plan figure to confirm, which is about BDT 98,000 a month if always on, and it sends your data out of the country. That is why rented cards serve only our lab.
FAQ
Questions we are asked
Can we use a server we already own?
Possibly. It needs a card with enough memory for the model the test bench chooses, mirrored drives and room for a UPS. We check it in the SI proof and say plainly if it will not do.
Do we need a generator?
Not for Team or Department, if a clean shutdown during a long cut is acceptable; the UPS handles that. Company is planned with a generator because 1,000 users expect the service to stay up.
Who fixes the card when it fails?
Your supplier, under the on-site warranty. We restore the service once the hardware is back, and the index and the settings come from your own backups.
Get the specification before the quote
The SI proof sizes the server on your documents and your users, so your supplier quotes the right box. Three weeks, BDT 5.5 lakh.