Muhammad Dhiyaul AthaThis is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass. ...
This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass.
EcoID is an offline-first plant identification tool for field observations.
The idea is simple: take a laptop or phone outside, photograph a real plant, and use a local AI model to get identification candidates. Then look at the plant yourself and verify the result.
EcoID is designed for:
The current model supports five plant classes:
Mangifera indica
Cocos nucifera
Musa acuminata
Carica papaya
Manihot esculenta
The AI does not present its output as a fact. It returns ranked identification candidates with confidence scores. The user then chooses whether the result is:
This human-in-the-loop flow matters because a model prediction is only a suggestion until somebody looks at the actual plant.
EcoID stores field observations locally, including the original photo, model prediction, alternative predictions, confidence score, verification status, notes, optional GPS coordinates, capture time from EXIF metadata, and user corrections.
The project is designed to make the computer useful for a short moment, then get the person back to observing the real world.
EcoID runs locally on a laptop or a phone-accessible local network. There is no public hosted demo because local execution is part of the project's privacy and offline design.
The trained model and downloaded dataset are not committed to the repository. Model weights and image data are kept outside Git because of their size and individual dataset licensing terms.
After training or obtaining the ONNX model locally:
python -m app --model models/ecoid-global-20261010.onnx --open
The app opens at:
http://127.0.0.1:8765
The main flow is:
Capture or upload a photo
↓
Run local inference
↓
Review possible identifications
↓
Verify, reject, or mark uncertain
↓
Save the field observation
↓
Review observations in history and on the local map
The source code is available on GitHub:
github.com/Bangkah/EcoID/tree/main
The source code is released under the MIT License. Dataset images retain their original licenses and attribution requirements.
The main parts of the project are:
app/ai/ — preprocessing, inference, model contracts, and result post-processingapp/observation/ — drafts, verification, EXIF metadata, maps, statistics, and exportsapp/storage/ — local SQLite persistence and image storageapp/ui/ — the local HTTP server and browser interfacescripts/ — dataset preparation, training, export, evaluation, and benchmarkingtests/ — unit, integration, offline, and browser testsThe architecture is:
Every image follows the same preprocessing path:
224x224
float32 tensor with shape (1, 3, 224, 224)
Keeping preprocessing in one shared module helps avoid train/serve skew between the training pipeline and the production inference path.
EcoID uses a MobileNetV3-Small image classifier exported to ONNX and executed with ONNX Runtime on the CPU.
The UI and storage layers depend only on a small model contract, so the inference backend can be replaced without rewriting the rest of the application.
The output is converted into a Top-3 list of candidates. If the confidence score is below the configured threshold, the result is marked as LOW_CONFIDENCE.
The initial threshold was calibrated using validation data. The current configuration is:
threshold = 0.73
On the current validation set, this produced:
These numbers are provisional and should be recalibrated when the field dataset grows.
The first real dataset download produced 2,447 usable licensed images, approximately 370 MB in total. The dataset was collected from iNaturalist using CC0, CC BY, and CC BY-SA images, with attribution metadata preserved.
The validated dataset contains approximately 2,000 training images across the five target classes, 300 validation images, and 149 negative samples across lookalike, non-plant, and low-quality groups.
The dataset pipeline includes:
The initial MobileNetV3-Small training run achieved:
On a separate 50-image evaluation set (10 images per class), the model achieved:
0.73: 94.1%
This evaluation set was collected from licensed iNaturalist images and is useful as a reproducible benchmark. It is not a substitute for a field evaluation with locally captured phone photos.
The ONNX export was verified against the PyTorch model:
1.9e-6
20/20 samplesThe repository also generates contact sheets for mandatory manual label review. Automated dataset validation passed with zero errors, but a complete human review of all contact sheets remains a dataset-quality task.
A prediction is first stored as a temporary draft. It becomes a permanent observation only after the user makes a verification decision.
This prevents an abandoned or uncertain prediction from silently becoming part of the observation history.
The flow is:
photo → identification draft → human verification → saved observation
Draft photos that are abandoned are automatically removed after 24 hours.
The application stores everything locally:
~/.ecoid/
├── ecoid.db
├── images/
├── thumbs/
└── drafts/
The application does not require a cloud API for inference. The browser UI also uses a strict Content Security Policy and does not load external scripts, fonts, or images.
GPS is opt-in. A user can attach a photo's EXIF location, use device geolocation, or enter coordinates manually.
The model was benchmarked on a Windows laptop using CPU inference:
| Stage | Average |
|---|---|
| Decode and resize | 6.3 ms |
| ONNX inference | 5.8 ms |
| Softmax and Top-K | 0.4 ms |
| Total | 12.4 ms |
The measured p95 total latency was 33.8 ms per image, below the project's target of 10 seconds per image.
The project includes:
The current local test result is:
222 tests passed
25 browser tests skipped unless browser testing is enabled
The CI workflow is defined in .github/workflows/ci.yml.
Plant identification is a good example of why open innovation matters.
A closed cloud API could return a label quickly, but it would also introduce several limitations:
Using open-source tools and local inference makes a different design possible. The complete pipeline can be inspected and adapted:
Open innovation also makes the project easier to learn from. Someone can inspect the training scripts, run the synthetic smoke test, replace the model, add new plant classes, or adapt the observation workflow for another field research problem.
The goal is not to pretend that a small classifier can replace botanical expertise. The goal is to build a transparent tool that helps people notice, record, and learn from the plants around them.
This is an initial working model, not a finished botanical identification system.
The most important remaining limitation is domain coverage. The 50-image independent evaluation set is sourced from licensed iNaturalist images, while a dedicated evaluation set of locally captured Indonesian field photos is still needed.
The next steps are:
This project was developed with an AI coding assistant using Copilot SDK in VS Code.
The agent helped inspect the repository, understand the architecture, connect the AI pipeline to the local observation workflow, run the dataset pipeline, train the initial model, verify ONNX parity, calibrate the confidence threshold, and validate the test suite.
The agent session is optional for this submission and is not linked here.
I am entering the overall Hacktoberfest Open-Source AI Challenge.
I am not entering a partner category because EcoID does not currently use a partner-specific technology.