Copy-Move and Splicing Detection in Images
Copy-move forgery duplicates a region within a single image to conceal or replicate content, while splicing composites elements from two or more source images. This topic covers the principal detection approaches: block-matching algorithms, keypoint-based descriptors, and deep-learning classifiers.
Copy-move and splicing are the two most prevalent forms of passive image forgery. In a copy-move attack, the forger selects a region within a single image, duplicates it, and pastes it over another area of the same image, typically to conceal an object, a person, or a marking. In a splicing attack, content from at least one external source image is composited into a target image.
Both operations alter the image's internal statistical structure in detectable ways. Copy-move leaves behind self-similar regions whose pixel statistics, noise patterns, and compression artifacts match more closely than any two genuine regions of that image would. Splicing introduces boundary discontinuities in noise level, colour channel statistics, JPEG blocking, and camera sensor fingerprint. Detection methods exploit these traces to localise the manipulated area and, in some cases, identify the source of the spliced content.
The forensic need for these detection methods is driven by their use in criminal and civil proceedings. Fabricated evidence photographs, altered surveillance stills, and manipulated news imagery have appeared in courts across multiple jurisdictions. A skilled examiner must apply the appropriate detection pipeline to the questioned image, document the methodology, and communicate the findings to a non-specialist fact-finder. No single algorithm is conclusive: convergent output from multiple independent methods, each flagging the same region, is the standard of practice.
Research in this area accelerated sharply after 2004, when widely used image editors became accessible enough for non-experts to produce plausible fakes. Detection methods have since evolved through three broad generations: hand-crafted block-matching algorithms, local keypoint descriptors, and, from around 2015 onward, convolutional neural networks.
Each generation inherited the limitations of the previous one and introduced its own. Understanding all three is necessary for a practitioner, because operational images arrive with unknown editing histories, and no single generation of methods is universally reliable.
By the end of this topic you will be able to:
- Distinguish copy-move from splicing forgery and describe the distinct artifact signature each leaves in the image.
- Explain how block-matching algorithms detect copy-move regions, including their key failure modes under geometric transformation and JPEG re-compression.
- Describe how SIFT and related keypoint descriptors are applied to forgery detection and what advantages they offer over block-matching.
- Summarise how deep-learning models are trained and tested for forgery detection and identify the generalisation risks that limit their use in casework.
- Outline the documentation and testimony requirements for presenting copy-move or splicing findings as evidence in court.
- Copy-move forgery
- An intra-image manipulation in which a patch copied from one location within the image is pasted over another location in the same image. The forged region and the source region share identical or near-identical statistical properties, which block-matching and keypoint methods exploit for detection.
- Splicing
- An inter-image manipulation in which content from one or more external images is inserted into a target image. Splicing introduces cross-boundary inconsistencies in noise level, colour statistics, JPEG blocking grid alignment, and camera fingerprint (PRNU).
- Block-matching
- A copy-move detection strategy that divides the image into overlapping fixed-size blocks, computes a compact feature vector per block, sorts vectors, and identifies suspiciously similar block pairs. The spatial relationship between matched pairs reveals the direction and magnitude of the clone operation.
- SIFT (Scale-Invariant Feature Transform)
- A keypoint descriptor algorithm that detects interest points and computes 128-dimensional descriptors invariant to scale and rotation. In forgery detection, matching SIFT descriptors within a single image identifies regions with identical local structure that would not be present in an authentic photograph.
- Passive forgery detection
- Detection that operates on the image data alone, without any pre-embedded watermark or signature. Distinguished from active methods such as digital watermarking and C2PA provenance signing, which require the capture device or publishing pipeline to embed authentication data at the time of creation.
- Localisation map
- The output of a forgery detection algorithm that marks, at pixel or block resolution, which regions of the image are identified as manipulated. A binary or heatmap localisation output is the primary forensic deliverable: it tells the fact-finder where the forgery occurred, not merely that one occurred.
Why forgeries leave traces
Every digital image is a statistical object. Its pixel values, their spatial correlations, their frequency-domain distribution after compression, and the pattern of sensor noise they carry all reflect the specific physics of the capture device and the specific sequence of processing operations applied. Forgery breaks one or more of these statistical regularities.
For copy-move, the break is internal consistency. A genuine photograph taken by a single camera has one noise level across the frame, one demosaicing pattern, one focus blur distribution, and one JPEG compression grid.
When a patch is copied and pasted, the pasted area and the original area have statistically identical texture and noise, a coincidence that is essentially impossible in a genuine image. If the forger applies geometric transformations such as rotation or scaling to disguise the copied region, the transformation itself leaves traces: interpolation artifacts, resampling periodicity, and a changed local noise spectrum.
For splicing, the break is cross-region consistency. A composited image contains content from cameras with different noise characteristics, different colour matrices, different JPEG quality settings, and potentially different acquisition conditions (lighting angle, colour temperature).
The boundary between the spliced region and the background may show an abrupt change in noise variance, a misalignment of the JPEG discrete cosine transform (DCT) blocking grid, or a sharp transition in the sensor's photo-response non-uniformity (PRNU) pattern. Detection methods target each of these inconsistency types.
Block-matching algorithms
Block-matching is the classical approach to copy-move detection, introduced in the work of Fridrich et al. in 2003. The algorithm divides the image into overlapping blocks of fixed size, typically 16 by 16 pixels, and represents each block as a compact feature vector.
DCT coefficient vectors and PCA-projected pixel vectors are the most common representations. The full set of feature vectors is then sorted lexicographically, and adjacent entries in the sorted list are compared. Pairs of blocks whose feature vectors fall within a small distance threshold are flagged as candidate copies.
From the candidate pairs, the algorithm computes offset vectors: each pair has a spatial offset equal to the vector from one block's position to the other's. Genuine block pairs that arise from chance similarity are scattered across many offset values. Cloned regions produce clusters of pairs sharing a consistent offset, because every block in the source region and its copy is displaced by the same translation. Post-processing uses this clustering to suppress false matches and produce a binary localisation map.
| Feature representation | Strengths | Weaknesses |
|---|---|---|
| Raw pixel blocks | Simple, no transformation loss | Sensitive to JPEG re-compression and minor brightness change |
| DCT coefficients | Compact, matches JPEG structure | Computationally expensive for large images; affected by rotated copies |
| PCA projections | Dimensionality-reduced, fast sort | Requires training set to compute PCA basis; less interpretable |
| Zernike moments | Rotation-invariant | Higher false-positive rate on textured regions; slow to compute |
Block-matching has two well-known failure modes. First, geometric transformation: if the forger rotates, scales, or shears the copied region before pasting, the feature vectors of corresponding blocks no longer match under direct comparison.
Extended variants using rotation-invariant features (Zernike moments, log-polar transforms) address this at a computational cost. Second, JPEG re-compression: saving the forged image as JPEG alters pixel values throughout, including in the copied region. The DCT rounding introduced by compression can push block pairs that were identical before save below threshold. Using a looser similarity threshold recovers some of these pairs but increases false positives.
Keypoint-based detection
Keypoint-based methods apply local feature descriptors developed originally for object recognition and image stitching to the problem of finding self-similar regions within a single image. SIFT (Scale-Invariant Feature Transform) detects interest points at multiple image scales and computes 128-dimensional gradient-based descriptors that are invariant to scale, rotation, and moderate illumination change. SURF and ORB (Oriented FAST and Rotated BRIEF) are faster variants with similar intra-image matching capability.
To detect copy-move with SIFT, the examiner extracts keypoints from the entire image and matches descriptors within the same image. In a genuine photograph, matched pairs arise from natural textures (brick walls, fabric) that have repeated structure, but the geometric relationship between matched pairs is inconsistent: they point in many directions and distances.
In a copy-move forgery, the matched pairs involving the cloned region produce a geometrically consistent cluster, because all the copies are displaced by the same affine transformation from their source. Grouping matches by homography consistency, for example using RANSAC (Random Sample Consensus), isolates the forgery cluster from background matches.
Keypoint methods handle rotated and scaled copies well. SIFT descriptors are computed relative to the dominant local gradient orientation, so a region copied and rotated 45 degrees still produces descriptors that match the source region's descriptors.
This makes keypoint methods considerably more resilient than basic DCT block-matching against geometric post-processing. Their limitation is density: keypoints concentrate on edges and textures, so a copied region that falls in a smooth, featureless area (sky, a painted wall) generates few or no keypoints and may be missed entirely. Block-matching covers featureless regions but fails on geometric transformation; combining both methods mitigates the blind spots of each.
Deep-learning approaches
From around 2015, convolutional neural networks began outperforming hand-crafted methods on standard forgery benchmarks. The early approach trained a CNN end-to-end on image patches labelled authentic or forged.
Subsequent architectures improved generalisation by feeding the network a noise residual (the image minus a denoised version, computed with a fixed filter such as a constrained convolutional layer) rather than raw pixels. Noise residuals amplify manipulation artifacts while suppressing semantic image content, making the network sensitive to the statistical traces of editing rather than to object appearance.
More recent models use dual-stream or frequency-domain inputs. One stream processes the RGB image; the other processes DCT coefficients or a high-pass filtered version. Fusion layers combine both streams to classify each image region. Transformer-based architectures have been applied to detect long-range inconsistencies that local convolutional filters miss. In controlled benchmark evaluations, these models achieve pixel-level localisation AUC above 0.95 on standard datasets such as COVERAGE, CASIA, and Columbia Uncompressed.
Splicing detection with deep learning follows the same structure but targets boundary artifacts rather than internal self-similarity. Networks learn to flag transitions in noise texture, lighting consistency, and colour statistics at region boundaries. The approach is sensitive to sophisticated compositing that blends these statistics at the boundary, and practitioners must be cautious about false negatives when the forger has used a feathered edge, gradient blend, or colour-matching tool.
Practical examination workflow
A casework examination begins with the questioned image in its originally received form. Any processing applied to the evidence file must be documented and reversible. The examiner first runs file-level checks: EXIF metadata, JPEG quantisation table, thumbnail-main discrepancy, and error-level analysis. These preliminary checks establish the image's processing history and may reveal obvious manipulation before algorithm-level detection begins.
Detection is then applied in layers. Block-matching is fast and covers featureless regions; keypoint matching covers geometric transformations; noise residual analysis targets splicing boundaries; PRNU analysis (where a reference noise pattern from the claimed source camera is available) tests device attribution.
Each method produces a localisation map. Convergence across maps strengthens the finding: if three independent methods flag the same region, the probability of all three producing a false positive simultaneously is very low. Divergence between maps indicates either a method-specific false positive or a complex forgery combining multiple techniques.
The examination report documents: the received file hash (SHA-256), the software and algorithm versions used, the parameter settings (block size, threshold, keypoint detector version), the output localisation maps, and a plain-language interpretation. Where the examiner concludes that manipulation is detected, the report quantifies confidence and lists alternative explanations considered and ruled out. Where results are inconclusive, this must be stated explicitly.
Legal standards and court presentation
Image authentication findings are admitted as expert evidence under jurisdiction-specific rules. In the United States, the Daubert standard requires that expert testimony be grounded in a testable, peer-reviewed methodology with a known or estimable error rate. In England and Wales, the Law Commission's report on expert evidence and the Criminal Practice Directions require disclosure of the basis and methodology.
In India, the Bharatiya Sakshya Adhiniyam 2023 (BSA 2023, which replaced the Indian Evidence Act 1872) governs the admissibility of electronic evidence and requires that digital records be certified as authentic by a person responsible for the device or process. Across the European Union, the eIDAS regulation and national procedural codes create comparable admissibility frameworks for electronic evidence.
Common challenges during testimony include: the opposing expert claiming the detection software is a black box with unknown false-positive rates; the suggestion that JPEG compression artefacts account for the flagged region; and the argument that the detection method has not been validated on images of this type. The examiner answers these by citing peer-reviewed validation studies, providing the software's published false-positive benchmarks, and describing the convergence of multiple independent methods in the specific examination.
International cases, such as those before the International Criminal Court (ICC) or the European Court of Human Rights, require the examiner to explain the methodology without assuming familiarity with any particular national evidentiary framework. The ICC has received forensic image authentication evidence in several proceedings, and the practice direction is to present the method's scientific basis, the output, and its limitations in terms that a tribunal of any legal tradition can evaluate.
What distinguishes copy-move from splicing forgery at the artifact level?
Key Takeaways
- Copy-move forgery duplicates a region within a single image and leaves internally self-similar regions; splicing composites content from external sources and leaves cross-region inconsistencies in noise, colour, and sensor fingerprint.
- Block-matching algorithms detect copy-move by finding block pairs with suspiciously similar feature vectors and consistent spatial offsets, but fail when the copied region has been geometrically transformed or heavily JPEG re-compressed.
- SIFT and related keypoint descriptors are scale and rotation invariant, making them more resilient than block-matching against geometric post-processing, but they miss copies in featureless regions where few keypoints are detected.
- Deep-learning models achieve high benchmark accuracy but degrade on forgeries produced by methods outside their training distribution; in casework, their output must be corroborated by classical methods and the training data must be documented.
- Court-ready findings require documented chain of custody, disclosed software versions and parameter settings, and convergent localisation from multiple independent methods, with explicit acknowledgement of each method's limitations and error rates.
What is the difference between copy-move and splicing forgery?
How do block-matching algorithms detect copy-move forgery?
What are keypoint-based detectors and why are they useful for tamper detection?
Can deep learning reliably detect image forgeries?
How is copy-move or splicing evidence presented in court?
Test yourself on Multimedia Authentication and Deepfake Forensics with free, timed mocks.
Practice Multimedia Authentication and Deepfake Forensics questions