Vector Logic · round 0 review pack

Papparti — what we propose to build

An on-demand services marketplace with a voice-first interface: a customer speaks a request in Hindi or Hinglish, and the system understands it, prices it and books it. This pack is the estimate, the scope, and the demos to judge it by.

Print-ready PDFs — one document to read, or five

Full project report — 90 pages, everything

One file containing the lot, with a contents page: executive summary and how to read it, quotation, proposal, cost estimate, the GPU decision, commercial terms and payment schedule, and appendices listing every figure and its source. If you read one thing, read this. Each part still carries its own cover and its own contents, so a section pulled out and sent onward on its own still stands up.

1.31 MB · 90 pages · A4 portrait · page numbers on every page

The same material as five separate PDFs

For when a finance reviewer should not have to page past the proposal, or an adviser wants only the estimate. Every number, every page number and every table of contents is identical to the combined report — these are the same documents, not a separate edit.

These are review copies, not signature copies. Fields only Papparti can confirm — registered entity name, signatory and designation, billing address, GSTIN, and the city your customers actually speak from — are marked as clearly-labelled placeholders rather than filled with a guess. The front-matter PDF lists every one of them on a single page. Send those details and we will issue the signed version.

Read first

Proposal — the full case

Ten sections: what you are buying, why a voice interface, the three models in parameter terms, the Phase 0 feasibility benchmark, the phase-by-phase scope, commercials, handover, risks stated plainly, what we need from you, and every assumption.

38 KB · markdown

GPU: rent or buy? — the running-cost decision

The monthly figure is the one the client pays forever, so it deserves its own case. A shown break-even, live-researched hardware and colocation prices each with a source and read date, the case against buying argued as hard as the case for it, and the exact condition that would flip the answer.

24 KB · markdown · recommendation: rent

Cost estimate — the numbers behind it

Every figure traced to a live vendor price or an explicitly-labelled assumption, with the source cited. Includes the finding that the original budget expectation was roughly one third of the real cost, and why.

24 KB · markdown

On the one number that matters most: the one-time cost is ₹78,72,318 excluding GST (₹92,89,335 including 18%), plus a monthly ₹1,26,108 after handover for hosting and support. On handover the models become yours. The proposal explains why this is not the ₹20–30 lakh band you might have been expecting, and what a smaller Phase 0-only engagement would cost.

Reading the architecture diagram on a phone. The diagram is drawn at a fixed metric so no label is ever shrunk or clipped. On a narrow screen it sits in a horizontal rail you can swipe, and a plain-text list of the same six steps in words appears beneath it. Nothing is cut off — if a step seems missing on a small screen, swipe the diagram or read the text list.

The talking avatar — a full Hindi conversation, out loud

This is the one artefact that shows the product rather than describing it: the real vector rig, animating, with the real three-turn Hindi conversation on it — a caller and the assistant, six clips, two voices. Fifty-one seconds, the full dialogue. Watch the mouth and listen at the same time.

Measured on this clipValue
Duration51.0 s — the full six-turn dialogue
Streamsvideo + audio
Pixel change, rest vs speaking5.7% of the frame
Frame rate24 fps

Three things this is, and two it is not. It is the real rig, the real voice, and real motion — the mouth opens on a syllable rhythm and an independent vision pass reads the rest state as “a thin, closed line” and the speaking state as “open, visible dark interior”.

It is not phoneme-accurate lip sync: the mouth is driven by a speech envelope, not by forced alignment to the audio's actual phonemes, so it is in time with the speech rather than shaped by it. And it is not a recording of a person — both voices are synthesised. We would rather show you an honest in-time mouth than a demo that implies perfect lip sync and delivers a surprise in Phase 1.

Motion study — no audio

The same rig with the motion isolated and no voice, so you can judge the animation itself.

The talking avatar — two rendering tiers

One character, drawn so it animates on the phone itself: no GPU, no server round-trip per frame, no per-conversation cost. Two tiers so a budget phone gets a simpler face, not a degraded one.

High-configuration tier, resting
High-configuration — resting Shaded skin, modelled eye sockets, visible ears, tapered neck.
High-configuration tier, mid-utterance
High-configuration — speaking The same face mid-utterance; the mouth region redraws 35.8% between frames.
Low-configuration tier, resting
Low-configuration — resting Deliberately flat: one skin colour, no gradients, block neck.
Low-configuration tier, mid-utterance
Low-configuration — speaking The same character, reduced, so a cheap phone still animates smoothly.
Measured propertyHigh tierLow tier
Tonal steps across the cheek30 values1 value
Total tonal range131 of 2550
Earsmodelledomitted by design
Neck51% cubic taperstraight block
Horizontal overflow at 320–1920 px0 px0 px

Open the live rig The actual vector file these images came from.

What these are not. Not a photoreal likeness, and not finished art direction. A photoreal animated face cannot run on a mid-range phone at zero marginal cost, so we are not showing you a photoreal demo and then delivering a rig. No human has judged the art direction yet — that decision is yours, and these four images are its input.

Voice samples — listen to these two first

Real output from a compact CPU-only engine on our own hardware. It is the fallback path, not the final brand voice — that is recorded in a studio session with a chosen speaker, and it is a priced line item.

A sentence carrying a price and a duration — male

The hardest case, and the one that carries money risk: a mis-spoken numeral is a wrong quote. Listen to this one first.

A sentence carrying a price and a duration — female

The same sentence, the other voice, so you can compare registers.

Short Hindi greeting — male

Naturalness of the opening.

Short Hindi greeting — female

The same greeting, the other voice.

Our recommendation: the male voice for the recording session — measured basis, not taste. On the price-bearing sentence it sits further from clipping (−16.2 dBFS against −12.8 dBFS). Your ear on your own script decides; that is what the studio session is for.

How every file was produced and measured →  ·  see the waveform (the audio is real, not a silent file)

No listening test has been run and no MOS exists. We measured duration, peak level, RMS and non-silent proportion objectively — but no human has rated this voice, and we quote no quality score from anywhere.

Honest status