Skip to content

Copy-Move and Splicing Detection in Images

Copy-move forgery duplicates a region within a single image to conceal or replicate content, while splicing composites elements from two or more source images. This topic covers the principal detection approaches: block-matching algorithms, keypoint-based descriptors, and deep-learning classifiers.

By Reviewed by Sourabh

Last updated:

Copy-move and splicing are the two most prevalent forms of passive image forgery. In a copy-move attack, the forger selects a region within a single image, duplicates it, and pastes it over another area of the same image, typically to conceal an object, a person, or a marking. In a splicing attack, content from at least one external source image is composited into a target image.

Both operations alter the image's internal statistical structure in detectable ways. Copy-move leaves behind self-similar regions whose pixel statistics, noise patterns, and compression artifacts match more closely than any two genuine regions of that image would. Splicing introduces boundary discontinuities in noise level, colour channel statistics, JPEG blocking, and camera sensor fingerprint. Detection methods exploit these traces to localise the manipulated area and, in some cases, identify the source of the spliced content.

The forensic need for these detection methods is driven by their use in criminal and civil proceedings. Fabricated evidence photographs, altered surveillance stills, and manipulated news imagery have appeared in courts across multiple jurisdictions. A skilled examiner must apply the appropriate detection pipeline to the questioned image, document the methodology, and communicate the findings to a non-specialist fact-finder. No single algorithm is conclusive: convergent output from multiple independent methods, each flagging the same region, is the standard of practice.

Research in this area accelerated sharply after 2004, when widely used image editors became accessible enough for non-experts to produce plausible fakes. Detection methods have since evolved through three broad generations: hand-crafted block-matching algorithms, local keypoint descriptors, and, from around 2015 onward, convolutional neural networks.

Each generation inherited the limitations of the previous one and introduced its own. Understanding all three is necessary for a practitioner, because operational images arrive with unknown editing histories, and no single generation of methods is universally reliable.

Detection MethodNo transformRotated or scaledcopyFeatureless regionJPEG re-compressedBlock-matching(DCT / PCA)DetectsFailsDetectsDegradesKeypoint / SIFTDetectsDetectsMissesPartialDeep learning(noise residual)DetectsDetectsDetectsDegrades if untrainedDetects reliablyFails or degradesMethod / scenario label
Block-matching detects featureless copies but fails on rotation; SIFT handles rotation but misses featureless regions; combining both closes each method's blind spot.

By the end of this topic you will be able to:

  • Distinguish copy-move from splicing forgery and describe the distinct artifact signature each leaves in the image.
  • Explain how block-matching algorithms detect copy-move regions, including their key failure modes under geometric transformation and JPEG re-compression.
  • Describe how SIFT and related keypoint descriptors are applied to forgery detection and what advantages they offer over block-matching.
  • Summarise how deep-learning models are trained and tested for forgery detection and identify the generalisation risks that limit their use in casework.
  • Outline the documentation and testimony requirements for presenting copy-move or splicing findings as evidence in court.
Key terms
Copy-move forgery
An intra-image manipulation in which a patch copied from one location within the image is pasted over another location in the same image. The forged region and the source region share identical or near-identical statistical properties, which block-matching and keypoint methods exploit for detection.
Splicing
An inter-image manipulation in which content from one or more external images is inserted into a target image. Splicing introduces cross-boundary inconsistencies in noise level, colour statistics, JPEG blocking grid alignment, and camera fingerprint (PRNU).
Block-matching
A copy-move detection strategy that divides the image into overlapping fixed-size blocks, computes a compact feature vector per block, sorts vectors, and identifies suspiciously similar block pairs. The spatial relationship between matched pairs reveals the direction and magnitude of the clone operation.
SIFT (Scale-Invariant Feature Transform)
A keypoint descriptor algorithm that detects interest points and computes 128-dimensional descriptors invariant to scale and rotation. In forgery detection, matching SIFT descriptors within a single image identifies regions with identical local structure that would not be present in an authentic photograph.
Passive forgery detection
Detection that operates on the image data alone, without any pre-embedded watermark or signature. Distinguished from active methods such as digital watermarking and C2PA provenance signing, which require the capture device or publishing pipeline to embed authentication data at the time of creation.
Localisation map
The output of a forgery detection algorithm that marks, at pixel or block resolution, which regions of the image are identified as manipulated. A binary or heatmap localisation output is the primary forensic deliverable: it tells the fact-finder where the forgery occurred, not merely that one occurred.

Why forgeries leave traces

Every digital image is a statistical object. Its pixel values, their spatial correlations, their frequency-domain distribution after compression, and the pattern of sensor noise they carry all reflect the specific physics of the capture device and the specific sequence of processing operations applied. Forgery breaks one or more of these statistical regularities.

For copy-move, the break is internal consistency. A genuine photograph taken by a single camera has one noise level across the frame, one demosaicing pattern, one focus blur distribution, and one JPEG compression grid.

When a patch is copied and pasted, the pasted area and the original area have statistically identical texture and noise, a coincidence that is essentially impossible in a genuine image. If the forger applies geometric transformations such as rotation or scaling to disguise the copied region, the transformation itself leaves traces: interpolation artifacts, resampling periodicity, and a changed local noise spectrum.

For splicing, the break is cross-region consistency. A composited image contains content from cameras with different noise characteristics, different colour matrices, different JPEG quality settings, and potentially different acquisition conditions (lighting angle, colour temperature).

The boundary between the spliced region and the background may show an abrupt change in noise variance, a misalignment of the JPEG discrete cosine transform (DCT) blocking grid, or a sharp transition in the sensor's photo-response non-uniformity (PRNU) pattern. Detection methods target each of these inconsistency types.

Block-matching algorithms

Block-matching is the classical approach to copy-move detection, introduced in the work of Fridrich et al. in 2003. The algorithm divides the image into overlapping blocks of fixed size, typically 16 by 16 pixels, and represents each block as a compact feature vector.

DCT coefficient vectors and PCA-projected pixel vectors are the most common representations. The full set of feature vectors is then sorted lexicographically, and adjacent entries in the sorted list are compared. Pairs of blocks whose feature vectors fall within a small distance threshold are flagged as candidate copies.

From the candidate pairs, the algorithm computes offset vectors: each pair has a spatial offset equal to the vector from one block's position to the other's. Genuine block pairs that arise from chance similarity are scattered across many offset values. Cloned regions produce clusters of pairs sharing a consistent offset, because every block in the source region and its copy is displaced by the same translation. Post-processing uses this clustering to suppress false matches and produce a binary localisation map.

Feature representationStrengthsWeaknesses
Raw pixel blocksSimple, no transformation lossSensitive to JPEG re-compression and minor brightness change
DCT coefficientsCompact, matches JPEG structureComputationally expensive for large images; affected by rotated copies
PCA projectionsDimensionality-reduced, fast sortRequires training set to compute PCA basis; less interpretable
Zernike momentsRotation-invariantHigher false-positive rate on textured regions; slow to compute

Block-matching has two well-known failure modes. First, geometric transformation: if the forger rotates, scales, or shears the copied region before pasting, the feature vectors of corresponding blocks no longer match under direct comparison.

Extended variants using rotation-invariant features (Zernike moments, log-polar transforms) address this at a computational cost. Second, JPEG re-compression: saving the forged image as JPEG alters pixel values throughout, including in the copied region. The DCT rounding introduced by compression can push block pairs that were identical before save below threshold. Using a looser similarity threshold recovers some of these pairs but increases false positives.

Keypoint-based detection

Keypoint-based methods apply local feature descriptors developed originally for object recognition and image stitching to the problem of finding self-similar regions within a single image. SIFT (Scale-Invariant Feature Transform) detects interest points at multiple image scales and computes 128-dimensional gradient-based descriptors that are invariant to scale, rotation, and moderate illumination change. SURF and ORB (Oriented FAST and Rotated BRIEF) are faster variants with similar intra-image matching capability.

To detect copy-move with SIFT, the examiner extracts keypoints from the entire image and matches descriptors within the same image. In a genuine photograph, matched pairs arise from natural textures (brick walls, fabric) that have repeated structure, but the geometric relationship between matched pairs is inconsistent: they point in many directions and distances.

In a copy-move forgery, the matched pairs involving the cloned region produce a geometrically consistent cluster, because all the copies are displaced by the same affine transformation from their source. Grouping matches by homography consistency, for example using RANSAC (Random Sample Consensus), isolates the forgery cluster from background matches.

Keypoint methods handle rotated and scaled copies well. SIFT descriptors are computed relative to the dominant local gradient orientation, so a region copied and rotated 45 degrees still produces descriptors that match the source region's descriptors.

This makes keypoint methods considerably more resilient than basic DCT block-matching against geometric post-processing. Their limitation is density: keypoints concentrate on edges and textures, so a copied region that falls in a smooth, featureless area (sky, a painted wall) generates few or no keypoints and may be missed entirely. Block-matching covers featureless regions but fails on geometric transformation; combining both methods mitigates the blind spots of each.

Deep-learning approaches

From around 2015, convolutional neural networks began outperforming hand-crafted methods on standard forgery benchmarks. The early approach trained a CNN end-to-end on image patches labelled authentic or forged.

Subsequent architectures improved generalisation by feeding the network a noise residual (the image minus a denoised version, computed with a fixed filter such as a constrained convolutional layer) rather than raw pixels. Noise residuals amplify manipulation artifacts while suppressing semantic image content, making the network sensitive to the statistical traces of editing rather than to object appearance.

More recent models use dual-stream or frequency-domain inputs. One stream processes the RGB image; the other processes DCT coefficients or a high-pass filtered version. Fusion layers combine both streams to classify each image region. Transformer-based architectures have been applied to detect long-range inconsistencies that local convolutional filters miss. In controlled benchmark evaluations, these models achieve pixel-level localisation AUC above 0.95 on standard datasets such as COVERAGE, CASIA, and Columbia Uncompressed.

Splicing detection with deep learning follows the same structure but targets boundary artifacts rather than internal self-similarity. Networks learn to flag transitions in noise texture, lighting consistency, and colour statistics at region boundaries. The approach is sensitive to sophisticated compositing that blends these statistics at the boundary, and practitioners must be cautious about false negatives when the forger has used a feathered edge, gradient blend, or colour-matching tool.

Practical examination workflow

A casework examination begins with the questioned image in its originally received form. Any processing applied to the evidence file must be documented and reversible. The examiner first runs file-level checks: EXIF metadata, JPEG quantisation table, thumbnail-main discrepancy, and error-level analysis. These preliminary checks establish the image's processing history and may reveal obvious manipulation before algorithm-level detection begins.

Detection is then applied in layers. Block-matching is fast and covers featureless regions; keypoint matching covers geometric transformations; noise residual analysis targets splicing boundaries; PRNU analysis (where a reference noise pattern from the claimed source camera is available) tests device attribution.

Each method produces a localisation map. Convergence across maps strengthens the finding: if three independent methods flag the same region, the probability of all three producing a false positive simultaneously is very low. Divergence between maps indicates either a method-specific false positive or a complex forgery combining multiple techniques.

The examination report documents: the received file hash (SHA-256), the software and algorithm versions used, the parameter settings (block size, threshold, keypoint detector version), the output localisation maps, and a plain-language interpretation. Where the examiner concludes that manipulation is detected, the report quantifies confidence and lists alternative explanations considered and ruled out. Where results are inconclusive, this must be stated explicitly.

Check your understanding
Question 1 of 4· 0 answered

What distinguishes copy-move from splicing forgery at the artifact level?

Key Takeaways

  • Copy-move forgery duplicates a region within a single image and leaves internally self-similar regions; splicing composites content from external sources and leaves cross-region inconsistencies in noise, colour, and sensor fingerprint.
  • Block-matching algorithms detect copy-move by finding block pairs with suspiciously similar feature vectors and consistent spatial offsets, but fail when the copied region has been geometrically transformed or heavily JPEG re-compressed.
  • SIFT and related keypoint descriptors are scale and rotation invariant, making them more resilient than block-matching against geometric post-processing, but they miss copies in featureless regions where few keypoints are detected.
  • Deep-learning models achieve high benchmark accuracy but degrade on forgeries produced by methods outside their training distribution; in casework, their output must be corroborated by classical methods and the training data must be documented.
  • Court-ready findings require documented chain of custody, disclosed software versions and parameter settings, and convergent localisation from multiple independent methods, with explicit acknowledgement of each method's limitations and error rates.
What is the difference between copy-move and splicing forgery?
Copy-move forgery copies a region from within the same image and pastes it elsewhere in that image, typically to conceal an object or duplicate a feature. Splicing composites content taken from two or more different source images into a single output image. Both are passive forgeries detectable through statistical analysis, but they leave different artifacts: copy-move leaves internal self-similarity, while splicing leaves cross-image boundary inconsistencies.
How do block-matching algorithms detect copy-move forgery?
Block-matching algorithms divide the image into overlapping fixed-size blocks, compute a feature vector for each block (DCT coefficients, PCA projections, or raw pixel values), sort the feature vectors lexicographically, and flag block pairs that are suspiciously similar. The flagged pairs form clusters whose spatial offset vectors reveal the direction and distance of the copy operation. Post-processing filters remove noise and highlight the cloned regions.
What are keypoint-based detectors and why are they useful for tamper detection?
Keypoint-based detectors such as SIFT, SURF, and ORB identify interest points in the image and compute descriptors that are invariant to scale, rotation, and moderate affine transformation. Matching descriptors across the image finds pairs of regions with visually identical local texture, even when the copied region has been rotated, scaled, or JPEG-compressed after pasting. This makes keypoint methods more resilient than block-matching against geometric post-processing of the forgery.
Can deep learning reliably detect image forgeries?
Convolutional neural networks trained on forgery datasets can detect both copy-move and splicing with high accuracy on benchmark images, and newer architectures using noise residuals or frequency-domain inputs generalise better than pixel-domain CNNs. However, all trained models degrade when tested on forgeries produced by methods not represented in training data. In casework, deep learning outputs are treated as a detection signal requiring validation with classical methods, not as a standalone conclusion.
How is copy-move or splicing evidence presented in court?
The examiner presents the detection method, the software or code used, validation data showing the method's error rates, and a localisation map showing which regions are flagged as manipulated. Courts in multiple jurisdictions, including under the Bharatiya Sakshya Adhiniyam 2023 in India, the Federal Rules of Evidence in the United States, and the Police and Criminal Evidence Act 1984 in England and Wales, require that digital evidence be accompanied by documentation of chain of custody and examiner qualification. The examiner must explain the method in terms accessible to a non-expert fact-finder and address the possibility of false positives.

Test yourself on Multimedia Authentication and Deepfake Forensics with free, timed mocks.

Practice Multimedia Authentication and Deepfake Forensics questions

Found this useful? Pass it along.

Share

Your journey to becoming a forensic professional starts here.

Practice with mock tests, learn from structured notes, and get your questions answered by a global forensic community, all in one place.