Tobacco barn detection from satellite imagery
Turns a GPS survey of known barns into a labelled training set, so a detector can look for the ones nobody has surveyed.
- Sole engineer
- 2026
- archived
- Python, YOLOv5, PyTorch, OpenCV
- GoPrime Systems, client project
- 58%
- 39%
- 0.42
- 99
Built for an employer against a client's survey data, so neither the imagery pipeline nor the coordinates are publishable.
Tobacco curing barns are small rectangular structures, and where they cluster says something useful about land use. Some areas have been surveyed on the ground, which produces a file of coordinates. Most have not, and the task is to find the barns nobody has walked to.
That is an object detection problem, and object detection needs labelled data. Normally someone draws a box around every instance in every image, which for something as small and repetitive as a barn means thousands of boxes before you learn anything.
The premise of this project was that the survey already contains the labels.
Labels from a survey rather than from a mouse
The ground survey arrives as KML: a placemark per structure, with a name and a latitude and longitude. The pipeline reads it, filters to the tobacco barns, and uses those coordinates for two things.
First, it decides where to look, generating a capture position per barn at a randomised zoom with up to fifteen metres of jitter on the camera position. The jitter is there so that not every training image is exactly centred on its barn. Without it the model can learn “the barn is in the middle of the frame” instead of what a barn looks like.
Second, it writes the labels from the same coordinates. For each captured image it computes the visible ground extent from the camera altitude and a field-of-view factor, derives metres per pixel, and projects every known barn into that image’s pixel space. Anything landing inside the frame becomes a YOLO bounding box. No human draws anything.
A useful side effect is that it labels barns the capture was not aimed at. If an image centred on one barn contains three others, all four get labelled, because the pipeline works from the full survey rather than from what an annotator happened to notice.
The projection is the difficult part
Flat-earth trigonometry puts a barn in roughly the right place, not the exact right place, because satellite imagery is not an orthographic projection. There is radial distortion and perspective foreshortening, and both grow with distance from the centre of the frame.
So the projection carries a correction that scales with radial distance from centre, pulling outlying objects back in, with a separate compression factor on the vertical axis. Box sizes get the same treatment: a base physical size in metres converted to pixels, scaled down slightly with distance from centre, then clamped so no box is absurd.
Those constants were arrived at by hand. There is a calibration routine in the code whose instructions amount to: generate the labels, look at them against the imagery, and if the boxes drift outward from centre reduce this number, if they drift inward increase it, if the sizes are all wrong adjust the field-of-view factor. That works, and it is also the main weakness of the run, for reasons below.
Two datasets that ask different questions
The training captures are barn-centred by construction: aim at a known barn, jitter slightly, capture. Nearly every image contains a barn near the middle.
Evaluating on data shaped like that would overstate the result, so a second generator tiles a region into a fixed grid at a constant zoom with fifteen percent overlap between neighbours, and labels those tiles from the same survey. A grid tile contains whatever it contains: no barns, one, or several.
That matches the real task, which is sweeping an area rather than pointing the camera at a known answer. Building the evaluation set to match the deployment condition rather than the training condition is the part of the method I would keep.
What the numbers say
After 150 epochs on 99 images, the detector reached 58% precision and 39% recall, with mean average precision at IoU 0.5 of 0.42, at the standard confidence threshold. Precision peaked around 63% mid-training and settled below that.
At 58% precision, two out of every five detections are not barns. At 39% recall, most real barns are missed. That is a prototype demonstrating a pipeline, not a system to point at unsurveyed land.
Why it fell short
Three causes, and they compound.
The dataset is tiny. Ninety-nine images is not a detection dataset. The pipeline can generate far more, since it is bounded only by survey size and capture throughput, but this run was not given more.
The label noise is systematic rather than random. This follows directly from calibrating by eye. Random label jitter mostly averages out during training. A projection tuned by hand carries a consistent bias, so every box in the set is wrong in the same direction by a similar amount, and the detector learns the offset as if it were signal.
Box sizes are assumed rather than measured. Every barn is labelled at the same base physical footprint, adjusted only for position in frame. Real barns vary, so the model is taught a size distribution that does not exist, and the IoU-based metrics penalise that. The gap between mAP at 0.5 and the much lower figure at stricter thresholds fits boxes that are roughly in the right place and the wrong shape.
The fix for all three is the same: hand-label a few hundred images properly, use them to measure how far off the projection is rather than eyeballing it, and use the corrected projection to generate the rest. The annotation tooling for that is in the repository and was set up but not carried through.
What I would keep
The idea that a GPS survey and an imagery source are together a labelled dataset generator, so a small amount of ground truth can be amplified into a large amount of training data. That transfers to anything where the objects of interest have been catalogued somewhere: infrastructure, water points, cell towers, informal settlements.
This run demonstrates the pipeline end to end, not the detector.