Repository navigation
Voiced/Unvoiced Consonant Issue w/ F0 Curve || CLI Renderer Issue #284
Description
Activity
I appreciate your discovery and research. Similar issues were also found by the devs, and here are my guesses and suggestions:
- Voicing curve can be improperly coupled with the pitch curve, especially in parts that should be unvoiced. DiffSinger didn't distinguish v/uv explicitly for convenience of architecture design and user experience. Although the model should be able to determine v/uv by phone sequences, some specific patterns in the pitch curve can still mislead the variance predictor.
- According to some other reports on this issue, I am working on some experiments on the Muon branch to see if this new optimizer is a partial cause. Are you using this branch?
- My main assumption is still that this is a dataset issue. Errors in variance curves can be due to some voiced parts accidentally labeled under unvoiced phones. Also I have to emphasize that many people use stuff like "trash phonemes" to label whatever they dislike, but this can cause unpredictable behavior and should be strictly avoided.
- You can do some experiments to "force" your model to learn silence. Add random silence paddings before and after your samples, as well as "SP" label. It's ugly but you can try.
- If you are on main branch, the default smooth width 0.12 may be too large. Try 0.06 which gives more precise curves.
- You really should at least turn off energy parameter. In the RMS view, energy is just breathiness + voicing. No one knows what the model will produce when the equation is broken.
Thank you for your prompt response!
I had no idea energy was breathiness and voicing. I also see the same recommendation mentioned in
BestPractices.md. I must have missed it! I'll be sure to re-read this as it must have slipped my mind.As for your other questions, I am using KakaruHayate's mix-LN, however I opted to use a scheduler without Muon optimization. I find that Muon in my specific case preforms slightly worse for my vocalist.
Furthermore, I do agree that this might be a dataset issue as well. I'll go back an revise some annotations to preform more tests and report back on this issue. As for trash phonemes, this model does not contain any. As for "SP" padding, I'll also give this a try!
Otherwise, the smooth width was actually set to 0.06! I had noticed the curves were not as precise a while back so I have had them already set to that.
Thank you so much for your insight!
Hello, I recently did some investigation into my dataset, and here's my progress that I would like to sync to you:
- The error mainly reflects in the voicing parameter predicted by the variance model. Disabling voicing will fix all or most of the errors. I reproduced what you described on my two speakers with voicing on, but they have no errors with voicing off.
- There was a bug in RMVPE that causes voiced parts unexpectedly marked as unvoiced. All 3 active branches (main/muon_lynxnet2/mix-ln) in this repo has fixed this bug. Maybe somehow related, but according to my experiments, this is not the root cause.
- There were indeed some errors in the data labels. I used a script to scan all the data, and excluded just several (10~20) pieces of data from my training set, and all the weird voiced parts were gone in the re-trained model. Note: this also proved that the unexpected voiced parts caused by label errors can spread through differnt speakers, so you have to scan and fix the whole dataset.
I uploaded the script to a new branch and here's how to use:
- Switch to
check-uv-voicingbranch. This branch is based on mix-ln branch, but we will only use its binarizer part. - Edit your acoustic model config. Enable voicing. Disable all augmentations. Clear all test_prefixes (no validation set).
- Run acoustic binarization.
- Edit check_ph_voicing.py, and set your input/output paths, phonemes to inspect (with their full name), and thresholds.
- Run the script, and you can see images generated, displaying possible data pieces with problems.
- Fix or exclude the pieces that you think to have real problems.
use_voicing_embed: true augmentation_args: random_pitch_shifting: enabled: false fixed_pitch_shifting: enabled: false random_time_stretching: enabled: false
- pinned this issue
on Jun 20, 2026

Hello! I hope this finds you well.
I've come across an issue recently and was wondering if they were related. I've recently trained a model with the variance parameters
Tension,Voicing,Energy, andBreathiness. It's a bit unorthodox as tension/voicing were successors to energy/breathiness, however I find that it gives me more freedoms with models.However, I've recently come across an issue where deep F0 curves can cause unvoiced consonants to become voiced. This causes issues as words/phrases become mispronounced. Using OpenUTAU's live curvature feature, I discovered that voicing would go upward rather than staying toward the bottom of the curve. While redrawing the curve fixes the solution, it makes using the model more tedious. I switched between using WORLD/VR for my hnsep and RMVPE/Parselmouth for my pe, however no combination solved this issue.
Without F0 modifications, the model retains [en/k] properly, however the [k] sound is enunciated as an unaspirated sound.

Here is an example of a tuned notes where [en/k] retains it's proper sound, however is now aspirated:

However in this example, the dip causes a slight increase in voicing here causes the [en/k] sound to transform into a [en/g] sound:

I've checked through the labels as well as run scripts to ensure there was not voicing within the consonant, yet the issue perssists.
While trying to figure this out I decided to use the CLI interface to see if it was an issue with OpenUTAU. I exported a sequence from OpenUTAU with the F0 curve maintained, and inferred variance using this command:
python scripts/infer.py variance "B:\Diffsinger\Blue.ds" --exp Singer1 --lang en --spk Spk1 --predict tension --predict energy --predict voicing --predict breathinessI proceeded to infer the acoustic part of the sequence with:
python scripts/infer.py acoustic "B:\Diffsinger\Blue.ds" --exp Singer1 --lang en --spk Spk1I'm not quite sure if there was an issue with my execution of the CLI command, so please let me know!
As shown in the image, the CLI inference has more errors, those of which seem related to voicing and breathiness specifically. Both were exported from the same checkpoints and the same steps number, except OpenUTAU was exported to ONNX before use as required. However, both yielded different results.
Would you be able to point me in the right direction of the CLI inference as well as config edits that can be made? I would love to be able to narrow down the issue regarding the voiced/unvoiced issue and CLI inference issue if related!