Vector Logic · round 0 review pack
An on-demand services marketplace with a voice-first interface: a customer speaks a request in Hindi or Hinglish, and the system understands it, prices it and books it. This pack is the estimate, the scope, and the demos to judge it by.
One file containing the lot, with a contents page: executive summary and how to read it, quotation, proposal, cost estimate, the GPU decision, commercial terms and payment schedule, and appendices listing every figure and its source. If you read one thing, read this. Each part still carries its own cover and its own contents, so a section pulled out and sent onward on its own still stands up.
For when a finance reviewer should not have to page past the proposal, or an adviser wants only the estimate. Every number, every page number and every table of contents is identical to the combined report — these are the same documents, not a separate edit.
These are review copies, not signature copies. Fields only Papparti can confirm — registered entity name, signatory and designation, billing address, GSTIN, and the city your customers actually speak from — are marked as clearly-labelled placeholders rather than filled with a guess. The front-matter PDF lists every one of them on a single page. Send those details and we will issue the signed version.
Ten sections: what you are buying, why a voice interface, the three models in parameter terms, the Phase 0 feasibility benchmark, the phase-by-phase scope, commercials, handover, risks stated plainly, what we need from you, and every assumption.
The monthly figure is the one the client pays forever, so it deserves its own case. A shown break-even, live-researched hardware and colocation prices each with a source and read date, the case against buying argued as hard as the case for it, and the exact condition that would flip the answer.
Every figure traced to a live vendor price or an explicitly-labelled assumption, with the source cited. Includes the finding that the original budget expectation was roughly one third of the real cost, and why.
On the one number that matters most: the one-time cost is ₹78,72,318 excluding GST (₹92,89,335 including 18%), plus a monthly ₹1,26,108 after handover for hosting and support. On handover the models become yours. The proposal explains why this is not the ₹20–30 lakh band you might have been expecting, and what a smaller Phase 0-only engagement would cost.
Reading the architecture diagram on a phone. The diagram is drawn at a fixed metric so no label is ever shrunk or clipped. On a narrow screen it sits in a horizontal rail you can swipe, and a plain-text list of the same six steps in words appears beneath it. Nothing is cut off — if a step seems missing on a small screen, swipe the diagram or read the text list.
This is the one artefact that shows the product rather than describing it: the real vector rig, animating, with the real three-turn Hindi conversation on it — a caller and the assistant, six clips, two voices. Fifty-one seconds, the full dialogue. Watch the mouth and listen at the same time.
| Measured on this clip | Value |
|---|---|
| Duration | 51.0 s — the full six-turn dialogue |
| Streams | video + audio |
| Pixel change, rest vs speaking | 5.7% of the frame |
| Frame rate | 24 fps |
Three things this is, and two it is not. It is the real rig, the real voice, and real motion — the mouth opens on a syllable rhythm and an independent vision pass reads the rest state as “a thin, closed line” and the speaking state as “open, visible dark interior”.
It is not phoneme-accurate lip sync: the mouth is driven by a speech envelope, not by forced alignment to the audio's actual phonemes, so it is in time with the speech rather than shaped by it. And it is not a recording of a person — both voices are synthesised. We would rather show you an honest in-time mouth than a demo that implies perfect lip sync and delivers a surprise in Phase 1.
The same rig with the motion isolated and no voice, so you can judge the animation itself.
One character, drawn so it animates on the phone itself: no GPU, no server round-trip per frame, no per-conversation cost. Two tiers so a budget phone gets a simpler face, not a degraded one.
| Measured property | High tier | Low tier |
|---|---|---|
| Tonal steps across the cheek | 30 values | 1 value |
| Total tonal range | 131 of 255 | 0 |
| Ears | modelled | omitted by design |
| Neck | 51% cubic taper | straight block |
| Horizontal overflow at 320–1920 px | 0 px | 0 px |
Open the live rig The actual vector file these images came from.
What these are not. Not a photoreal likeness, and not finished art direction. A photoreal animated face cannot run on a mid-range phone at zero marginal cost, so we are not showing you a photoreal demo and then delivering a rig. No human has judged the art direction yet — that decision is yours, and these four images are its input.
Real output from a compact CPU-only engine on our own hardware. It is the fallback path, not the final brand voice — that is recorded in a studio session with a chosen speaker, and it is a priced line item.
The hardest case, and the one that carries money risk: a mis-spoken numeral is a wrong quote. Listen to this one first.
The same sentence, the other voice, so you can compare registers.
Naturalness of the opening.
The same greeting, the other voice.
Our recommendation: the male voice for the recording session — measured basis, not taste. On the price-bearing sentence it sits further from clipping (−16.2 dBFS against −12.8 dBFS). Your ear on your own script decides; that is what the studio session is for.
How every file was produced and measured → · see the waveform (the audio is real, not a silent file)
No listening test has been run and no MOS exists. We measured duration, peak level, RMS and non-silent proportion objectively — but no human has rated this voice, and we quote no quality score from anywhere.