Belumind Science

Hard problems in AI and software — ours, and our clients'

We build and ship our own products, and we take on client work in AI and software engineering. Under both sits sustained work on systems and new technology, architecture, and AI infrastructure — written up in full, including the parts that are still open.

  • 5 products in production
  • 3 client engagements
  • 1 paper published

What we work on

The same engineering floor under our work and yours

Four disciplines, and none of them stops at the demo. Each one is in service of something already running — a product of ours or a client's system — or something we are actively measuring our way toward.

  • Product & client engineering

    Shipped software, ours and our clients'

    We take work all the way to something people use every day. Our own products — a converter built for analysts moving financial reports, an assistant grounded in Vietnamese law, a macOS utility that keeps everything on the machine — and client systems in AI and software where nothing off the shelf fits. Shipping is where design decisions get audited.

    • Web apps
    • macOS
    • Browser extensions
    • Client delivery
  • Systems & new technology

    Problems where the existing tools quit early

    ARMT schedules a single file transfer across every radio two devices share — channels separated by six orders of magnitude in bandwidth — and treats a dying link as a rebalance rather than an error. The measurement work behind that decision is the actual research.

    • Transport protocols
    • Multipath
    • Rust
    • Mobile radios
  • Architecture

    Layers with contracts, not diagrams that age

    Control planes, session lifecycles, and plugin boundaries left loose enough to absorb transports we have not built yet. Failure is modeled as a normal state rather than an exception path, which is most of what decides whether a system survives its second year.

    • Layered contracts
    • Plugin boundaries
    • Failure-first design
  • AI infrastructure

    Models small enough to check, pipelines that hold structure

    OCR and cell-by-cell table reconstruction that keeps a document's structure intact, retrieval grounded in primary legal sources, and speech-to-text with translation and summaries on top. Our published sign-language model is 0.74 MB and runs at 3.02 ms on CPU — small, checkable models are a position, not a constraint.

    • OCR pipelines
    • Grounded retrieval
    • Speech & translation
    • CPU inference

Client work

Bring us a problem

Alongside our own products, we take on client work in AI and software engineering — the problems where nothing off the shelf fits, where a model has to run somewhere awkward, or where the system has to keep working while parts of it fail. Below is some of what that has looked like.

  • A model that has to run on-device, on CPU, or on a budget that rules out an API call per request.
  • A document pipeline that has to preserve structure — tables, columns, meaning — not just extract text.
  • A system that has to degrade gracefully when a dependency, a network, or a device drops out.
  • An architecture review while changing it is still cheap, rather than after the second year.

Selected work

3 engagements
  • JICEEESG document processing

    Sustainability reporting that survives the paperwork

    ESG disclosure arrives as a mountain of documents — reports, certificates, and filings in every format a supplier chooses to send. We built the processing pipeline that turns that pile into structured, auditable data: extracting the figures that matter, preserving the tables and context they came from, and keeping every value traceable back to the page it was read from.

    • ESG
    • Document AI
    • Data extraction
    • Auditability
  • Vietnamese commercial bankeKYC identity verification

    Opening an account without walking into a branch

    Remote onboarding only works if the bank can be certain who is on the other end of the camera. We built the eKYC model behind that decision: reading Vietnamese ID documents, matching the face in front of the lens to the one on the card, and telling a live person apart from a photograph, a replay, or a generated face. It runs on whatever phone the customer already owns, in whatever lighting they happen to be in, and it has to be right — the cost of a wrong answer is a fraudulent account or a customer turned away.

    • eKYC
    • Liveness detection
    • Face matching
    • OCR
    • Fintech
  • GPRONoise detection model

    Teaching a model to hear what shouldn't be there

    Precision falls apart when a signal is buried in noise. We built the detection model that finds it — separating genuine signal from interference so downstream measurements can be trusted. Sharper inputs, fewer false readings, and a model tuned for the conditions it actually runs in rather than a clean benchmark.

    • Signal processing
    • Model training
    • Precision
    • Detection

Published in full

Research and engineering, in the open

Papers with the numbers attached, and write-ups that describe the problems before the answers exist. Both are how we think, not marketing for the products.

Latest paper

CSF: Contrastive Semantic Features for Direct Multilingual Sign Language Generation

A language-agnostic semantic framework enabling direct translation from any source language to sign language. Achieves 99.03% slot extraction accuracy across four languages with a lightweight 0.74 MB transformer running at 3.02ms inference on CPU.

Latest write-up

9 min read

A file transfer that refuses to die

We're building a transport protocol that treats every channel between two devices — every radio, every network it can reach — as one pool of capacity. When a channel fails mid-transfer, the transfer doesn't notice.

Start anywhere

The products, the papers, and the engineering write-ups all point at each other. Pick whichever one you would rather read first — or just get in touch.

or email hello@belumind.com