Towards a satellite image manipulation and deepfake localization benchmark dataset

Authors: Jacob Arndt, Debvrat Varshney, Philipe Dias, Nivedita Nukavarapu

Published: 2026-08-05 13:36:42+00:00

Comment: Accepted at IEEE IGARSS 2026

AI Summary

This paper addresses the critical need for high-quality, fine-grained manipulation datasets in remote sensing by introducing a prototype benchmark dataset for satellite image manipulation detection and localization. It offers 60 images (30 manipulated, 30 authentic) with ground-truth masks and metadata, created using copy-paste splicing and diffusion model inpainting, to support research in geospatial deepfake detection.

Abstract

Verifying the authenticity of satellite imagery has become increasingly critical given advances in generative artificial intelligence. Highly realistic synthetic imagery produced for malicious purposes (deepfakes) can have major consequences in the remote sensing domain, where this data is a fundamental source of information for science applications, planning, logistics, and monitoring. The remote sensing community lacks high-quality, fine-grained manipulation datasets suitable for training and evaluating detection and image forensics algorithms. Existing datasets are lacking and those that do exist either provide no ground truth masks for evaluating manipulation localization, or consist of entire images generated by GANs or diffusion models, which are inadequate for measuring localization performance. To address this gap, we describe a preliminary dataset construction process and prototype benchmark dataset for satellite image manipulation detection and localization. The dataset contains 60 images total, with 30 images carefully manipulated using three manipulation types including copy-paste splicing and diffusion model inpainting, and 30 authentic images. Each image is accompanied by a ground-truth mask and acquisition metadata, enabling both pixel-level localization metrics, image metadata studies, and analyses of how manipulation detection performance relates to image collection parameters. We describe the dataset construction process and present this initial release to support further research in image forensics and geospatial deepfake detection. The prototype dataset can be downloaded at https://huggingface.co/datasets/geodf/fmow-fake-small.


Key findings
The prototype dataset contains high-quality deepfakes with fewer visual artifacts compared to existing datasets, addressing a critical gap in the remote sensing community. It provides diverse image forgeries with ground truth masks and extensive metadata, making it suitable for evaluating deepfake localization methods. Despite its small size, it offers a more robust and realistic evaluation resource for remote sensing deepfake detection algorithms.
Approach
The authors construct a dataset by manipulating satellite images from the fMoW dataset using three techniques: simple copy-paste splicing, object copy-paste splicing, and diffusion model inpainting. Each manipulated image is accompanied by a ground-truth mask indicating the altered regions and associated acquisition metadata, enabling pixel-level localization metrics and analysis of detection performance.
Datasets
fMoW (Functional Map of the World)
Model(s)
Segment Anything Model (SAM), RSPaint (finetuned Stable Diffusion model)
Author countries
USA