CrySense AI Training Center
Train and evaluate an experimental audio classification model.
Pipeline
- 1Audio Dataset
- 2Preprocessing
- 3Feature Extraction
- 4Model Training
- 5Validation
- 6Model Export
- 7ESP32 Deployment
Dataset
Balance matters more than raw volume.
๐ผ Likely Hungry
Samples: 148
๐จ Possible Burp
Samples: 96
๐ฃ Possible Gas
Samples: 74
๐งท Possible Discomfort
Samples: 112
๐ด Likely Sleepy
Samples: 131
๐ Background / No Cry
Samples: 205
Labels are never fabricated automatically. Source, consent and license metadata stay attached to every sample.
Audio preprocessing
Sample rate
16000 Hz
Audio format
WAV
Channels
Mono
Audio window
1000 ms
Window overlap
500 ms
Bit depth
16-bit
Augmentation: background noise mixing, volume variation, small timing shifts.
Feature extraction
Raw audio โ spectrogram โ MFCC โ model input.
Configure AI model
TinyML Audio Classifier targeting ESP32-S3 constraints.
Model performance
Accuracy
82.4%
Precision
0.81
Recall
0.79
F1 score
0.80
Model size
180 KB
RAM required
120 KB
Inference time
75 ms
ESP32-S3
๐ข Compatible
Confusion matrix
Rows are actual classes, columns are predictions.
| HUN | BUR | GAS | DIS | SLP | |
|---|---|---|---|---|---|
| HUN | 42 | 2 | 1 | 3 | 2 |
| BUR | 3 | 28 | 4 | 2 | 1 |
| GAS | 2 | 5 | 24 | 4 | 1 |
| DIS | 4 | 2 | 3 | 31 | 3 |
| SLP | 2 | 1 | 1 | 3 | 38 |
Outlined cells mark confusable class pairs worth more samples (burp โ gas).
Prediction rules
Smoothing prevents single noisy frames from raising an alert.
Minimum confidence
70%
Consistency requirement
3 consecutive predictions
Cooldown between alerts
30 seconds
AI predictions are estimates and should always be interpreted together with context and normal caregiving judgment. CrySense AI is not a medical device.