CrySense AI Training Center

Train and evaluate an experimental audio classification model.

Pipeline

  1. 1Audio Dataset
  2. 2Preprocessing
  3. 3Feature Extraction
  4. 4Model Training
  5. 5Validation
  6. 6Model Export
  7. 7ESP32 Deployment

Dataset

Balance matters more than raw volume.

๐Ÿผ Likely Hungry

Samples: 148

๐Ÿ’จ Possible Burp

Samples: 96

Underrepresented class

๐Ÿ˜ฃ Possible Gas

Samples: 74

Underrepresented class

๐Ÿงท Possible Discomfort

Samples: 112

๐Ÿ˜ด Likely Sleepy

Samples: 131

๐Ÿ”‡ Background / No Cry

Samples: 205

Labels are never fabricated automatically. Source, consent and license metadata stay attached to every sample.

Audio preprocessing

Sample rate

16000 Hz

Audio format

WAV

Channels

Mono

Audio window

1000 ms

Window overlap

500 ms

Bit depth

16-bit

Augmentation: background noise mixing, volume variation, small timing shifts.

Feature extraction

Raw audio โ†’ spectrogram โ†’ MFCC โ†’ model input.

Configure AI model

TinyML Audio Classifier targeting ESP32-S3 constraints.

Model performance

Accuracy

82.4%

Precision

0.81

Recall

0.79

F1 score

0.80

Model size

180 KB

RAM required

120 KB

Inference time

75 ms

ESP32-S3

๐ŸŸข Compatible

Confusion matrix

Rows are actual classes, columns are predictions.

HUNBURGASDISSLP
HUN422132
BUR328421
GAS252441
DIS423313
SLP211338

Outlined cells mark confusable class pairs worth more samples (burp โ†” gas).

Prediction rules

Smoothing prevents single noisy frames from raising an alert.

Minimum confidence

70%

Consistency requirement

3 consecutive predictions

Cooldown between alerts

30 seconds

AI predictions are estimates and should always be interpreted together with context and normal caregiving judgment. CrySense AI is not a medical device.