A Weekend Challenge, a Colab Surprise, and an Honest 60%: Building AgriScout

Musa Badru

2026-09-28

How it started

I opened my Google AI Pro plan one day and noticed it now included Google Colab. That meant access to a real GPU without buying hardware. Around the same time, a colleague dared me:

Build an app over the weekend that can be as good as PlantVillage+, which is now unmaintained.

I took the bait. The result is AgriScout, an Android proof of concept for researchers and enumerators who need to scout cassava fields, often with poor or no connectivity. Everything runs on-device, offline.

Spoiler: it is not yet "as good as PlantVillage+". This post explains why, and why I think publishing that fact is more useful than hiding it.

The contract: five classes, fully offline

Before touching a model, I fixed the contract the app depends on:

Item Value
Classes CBB, CBSD, CGM (Cassava Green Mottle), CMD, Healthy
Model MobileNetV3-Small, ImageNet init
Input RGB, 224×224×3224 \times 224 \times 3224×224×3
Export Full INT8 TFLite / LiteRT
Runtime Android CPU interpreter, no network

A model that doesn't match this contract exactly does not ship. That rule was tested almost immediately.

Detour: the checkpoint I refused to ship

My first idea was to skip training and grab a pretrained Apache-2.0 cassava checkpoint. It converted cleanly, but into a six-class model with no documented label order. AgriScout's contract is a verified five-class one. Guessing which output index meant which disease would mean the app could confidently show the wrong diagnosis.

So I deleted the 50.8 MB source and its conversion output, so it couldn't be shipped by accident, and chose to train my own.

Lesson 1: a model with undocumented labels is not "almost usable". It is unusable.

First plan: train on my laptop (and why I didn't)

My original plan was a small CPU fine-tune on the Makerere five-class dataset (about 21k images at 224×224224 \times 224224×224). My estimates:

  • Train only the new head: roughly 1–3 hours
  • Fine-tune the whole network: roughly 4–12 hours, longer if the laptop throttles
  • INT8 conversion and checks: minutes

The integrated GPU isn't a practical TensorFlow accelerator, so that was slow and awkward. Colab's T4 changed the economics, which is where the Colab bundle paid off.

The pipeline

Each stage saves only its best checkpoint to Google Drive and is skipped on rerun if the checkpoint already exists. Colab runtimes disappear, so restart-safety matters more than it sounds.

Data, and the trap in the label names

The primary data is the Kaggle Cassava Leaf Disease Classification set:

Class Count
CBB 1,087
CBSD 2,189
CGM 2,386
CMD 13,158
Healthy 2,577

This is heavily imbalanced. I used inverse-frequency class weights:

wc=NK⋅ncw_c = \frac{N}{K \cdot n_c}wc​=K⋅nc​N​

where NNN is the number of images, K=5K = 5K=5 classes and ncn_cnc​ the count for class ccc. With the full counts, CBB gets w≈3.94w \approx 3.94w≈3.94 while CMD gets w≈0.33w \approx 0.33w≈0.33.

I also added a supplementary set, iCassava. Its directories are cbb, cbsd, cgm, cmd, healthy, and here is the trap: in iCassava, CGM means Cassava Green Mite. In the Kaggle data, CGM means Cassava Green Mottle. Same three letters, different diseases. Merging them would have quietly poisoned one class, with no error message to warn me. I excluded iCassava's cgm folder and kept only the four compatible labels.

Lesson 2: when combining datasets, never trust matching label strings. Check what they mean.

Getting a float16 model onto a phone

I trained with mixed_float16 to use the T4 efficiently. When I tried to convert that model directly to strict INT8, conversion failed because the float16 compute ops needed Select-TF-Ops (Flex) fallback, which I didn't want on Android.

The fix:

  1. Reload the best checkpoint.
  2. Build a fresh float32, inference-only MobileNetV3-Small.
  3. Copy the trained weights into it.
  4. Calibrate with 200 training images.
  5. Convert with TFLITE_BUILTINS_INT8 and INT8 input/output.

Quantization is affine: q=round⁡(x/s)+zq = \operatorname{round}(x / s) + zq=round(x/s)+z. For the input, s=1.0s = 1.0s=1.0 and z=−128z = -128z=−128, so raw pixel values 0..2550..2550..255 map to −128..127-128..127−128..127 with no normalization step. For the output, s=1/256s = 1/256s=1/256 and z=−128z = -128z=−128, so a probability is recovered as:

p=q+128256p = \frac{q + 128}{256}p=256q+128​

Lesson 3: train fast in mixed precision, but export from a clean float32 graph.

The honest results

Checkpoint Val. accuracy Macro F1
Fine-tuned 57.5% 0.511
Deeper fine-tune 59.0% 0.518
+ iCassava augmentation 60.2% 0.526

Macro F1 averages per-class F1 so that rare classes count equally:

Macro-F1=1K∑c=1K2 PcRcPc+Rc\text{Macro-F1} = \frac{1}{K}\sum_{c=1}^{K} \frac{2\,P_c R_c}{P_c + R_c}Macro-F1=K1​c=1∑K​Pc​+Rc​2Pc​Rc​​

Now the uncomfortable part. CMD makes up about 13158/21397≈61.5%13158 / 21397 \approx 61.5\%13158/21397≈61.5% of the data, so a model that answers "CMD" for everything scores about 61.5% accuracy. My best checkpoint scores 60.2%. On accuracy alone, it does not beat guessing the majority class. The macro F1 and the class weights are the reason it's still doing something useful on rarer classes, but I'm not going to dress this up.

On a small held-out iCassava subset (87 images) it scored 80.5% accuracy and 0.807 macro F1. That looks encouraging, but the 95% Wilson interval on 87 images at 80.5% is roughly 71% to 87%, and those images come from the same source as some of my training data. It is a sanity check, not a reliability claim.

What I'd fix next

  • Suspect the pipeline, not just the model. Scores like these usually point at something fixable. First on my list: verify preprocessing, since Keras's MobileNetV3 includes its own rescaling and normalizing twice is an easy mistake.
  • Measure the INT8 model separately. Report FP32 and INT8 accuracy side by side, and confirm the Android output matches Keras on identical images.
  • Split by farm, not by image. Random image splits leak near-duplicate photos from the same plant or field into validation, which flatters the numbers.
  • Collect local field data. Independently labelled Kenyan/Ugandan images with crop-pathology or extension partners, ideally 300–500 independent photos to start, evaluated per class.
  • Keep an unknown / other path for anything outside the five supported conditions.

Safety

AgriScout is an advisory proof of concept for researchers and enumerators, not a diagnostic tool. It should never give management recommendations from the top prediction alone, and it does not replace agronomist or laboratory confirmation. If the model file is missing, the app shows an "unavailable" state instead of inventing a result. If the model's output does not match the app's label contract, it refuses to run.

What I learned

  1. A weekend is enough to build the plumbing. It is nowhere near enough to earn trust.
  2. Undocumented labels and colliding label names are silent killers.
  3. Checkpoint, skip and resume everything when your compute is ephemeral.
  4. Report the baseline next to your accuracy. Otherwise the number means nothing.

So, colleague: challenge partly accepted. AgriScout runs offline on a phone, but "as good as PlantVillage+" will have to be earned with field data.

Tagged in: #kotlin#android#app#research#model-training

Subscribe to the newsletter

Get emails from me about Lorem ipsum dolor sit, amet consectetur adipisicing elit. Libero, ducimus..

5,432 subscribers including my Mom – 123 issues

Latest Posts

Search and see all posts