# Papparti — voice samples (round 0)

These are **real audio**, produced by a real text-to-speech engine running on our own
machine. They are not mock-ups, not stock clips, and not recordings of a person.

## What produced them

- **Engine:** a small CPU-only neural text-to-speech engine, run locally with no GPU.
- **Voices:** one male and one female Hindi voice.
- **Machine:** an Intel i3-6006U laptop, 4 threads, **no discrete GPU** — deliberately the
  weakest hardware we could test on.

## Why this engine, and not the production voice

This is a **compact fallback engine**, the kind of thing that runs when no GPU is available.
It is **not** the model that will ship as the brand voice — that is a fine-tuned model built
in Phase 0B, recorded with a chosen speaker in a studio session.

So read these samples as showing **the character of the fallback path**: what the product can
still say out loud on hardware with no accelerator at all. That is a genuine commercial
property — it means the voice keeps working when the GPU is busy — but it is **not** a
preview of the final brand voice, and we are not going to present it as one.

## Measured facts (measured here, not vendor claims)

| file | duration | peak | RMS | non-silent |
|---|---|---|---|---|
| `sample_hi_pratham.wav` (greeting) | 5.543 s | 32767 | -15.4 dBFS | 74.8% |
| `sample_hi_priyamvada.wav` (greeting) | 6.844 s | 32767 | -13.7 dBFS | 76.8% |
| `sample_num_pratham.wav` (with numbers) | 6.356 s | 32767 | -16.2 dBFS | 72.4% |
| `sample_num_priyamvada.wav` (with numbers) | 6.785 s | 32767 | -12.8 dBFS | 79.0% |

All measured with the Python standard-library `wave` module, reading the PCM samples
directly — **not** by trusting the engine's own log output.

**Real-time factor:** 0.13 on a short sentence, 0.24 on a longer one containing numerals. It
rises with text length, which is why two different figures for this engine appear in our
research (0.455) and here (0.13): they measured utterances of different lengths. Neither is
wrong; they are different workloads.

## The number test

`sample_num_*.wav` contains: *"आपकी बुकिंग कन्फर्म हो गई। प्लंबर दो बजे आएगा, कुल खर्च चार सौ
पचास रुपये होगा।"* — a booking confirmation carrying a time and a price.

This is the hardest thing a text-to-speech engine must get right for this product:
**a mis-spoken price is a money error.** Phase 0B carries it as a hard gate (number and
duration intelligibility must be 100%). The RMS on the numeric sentence is **lower** than on
the greeting (−16.2 vs −15.4 dBFS, male), which is the first evidence that numerals are
spoken with less energy — worth watching in the benchmark.

## Files

- `papparti-voice-*.mp3` — the greeting, normalised to −1 dBFS, 96 kbps mono
- `papparti-number-*.mp3` — the booking-with-price sentence
- `waveform-*.png` — waveform images; these exist so a reader can **see** that the audio is
  real audio and not a silent file
- `sample_hi_*.wav` — the two raw greeting renders at full quality

## Honest limitations

- This is a **compact fallback engine**, not the brand voice. Do not read it as the final
  voice quality.
- **The brand voice does not exist yet.** It requires a studio recording session before it
  can be fine-tuned.
- **No listening test has been run and no MOS (mean opinion score) exists.** Any quality
  score quoted from elsewhere is not a measurement of this product's voice.


---

## The conversation demo — `conversation/`

Two isolated sentences cannot show whether the product holds a **conversation**, which is the
whole product. This set is a **scripted three-turn exchange** between a caller and the assistant
— six clips, two voices, one role each, so the two speakers are distinguishable by ear.

**What is real and what is not, stated plainly:**

- **Real:** the audio, the measurements below, and the fact that this is a genuine end-to-end
  dialogue with a **time** and a **price** in it — the two items where a mistake becomes a money
  error, and the two a client should listen to most carefully.
- **Not real:** this is **not a live model output**. It is **not a recording of a person**. Both
  roles are **synthesised** — the customer's lines in a male voice, the assistant's in a female
  voice. The brand voice does not exist yet, because the studio session that produces it is a
  priced line in the estimate and has not been booked.

**Why both roles are synthesised rather than one being a human:** so the demo makes no claim it
cannot support. A half-human demo invites the client to read a scale that is not there.

| file | seconds | role |
|---|---:|---|
| `turn-01.mp3` | 6.19 | customer |
| `turn-02.mp3` | 9.59 | assistant |
| `turn-03.mp3` | 6.82 | customer |
| `turn-04.mp3` | 13.17 | assistant (contains the price) |
| `turn-05.mp3` | 4.86 | customer (confirms the price) |
| `turn-06.mp3` | 8.78 | assistant (confirms the slot) |
| `conversation-full.mp3` | 51.07 | the stitched take, 0.4 s between turns |

**The stitched take is also muxed onto the moving avatar** — see
`../avatar/avatar-speaking-conversation.mp4`. That clip is the one closest to the product: a
face that visibly speaks, with the voice audible, for the full dialogue.

**Honest limitation:** the mouth on that clip is driven by a **speech envelope**, not by forced
alignment to the phonemes in this audio. It is **in time with** the speech rather than
**shaped by** it. We would rather say so than let it be read as perfect lip sync.


---

### A footnote on the durations, in case anyone cross-checks with a media tool

Every figure in the table above is decoded from the audio samples themselves. A media tool
reading the MP3 **container** reports each file about **70 ms longer**, and that is expected:
MP3 adds a short silent padding frame at the start. It is why the stitched take is cited as
**50.999 s** (the turn durations plus five 0.4 s pauses, exactly) rather than the container's
51.069 s. Neither number is wrong; they measure the container and the sound respectively.
