Tai Vu, Robert Yang -- Stanford University
A comparative study of deep generative models for automatic anime sketch colorization. We implement and benchmark Neural Style Transfer, Conditional GAN (Pix2Pix), and CycleGAN on 17,769 sketch-color pairs, achieving state-of-the-art results of 220.5 FID and 0.76 SSIM with our modified C-GAN incorporating total variation regularization.
All models were evaluated on 100 held-out images using FID (lower is better) and SSIM (higher is better):
| Model | FID | SSIM (mean) | SSIM (std) |
|---|---|---|---|
| Neural Style Transfer | 345.506 | 0.6547 | 0.0989 |
| CycleGAN | 272.619 | 0.7238 | 0.0824 |
| C-GAN (Pix2Pix) | 227.948 | 0.7469 | 0.0741 |
| C-GAN + TV Loss (Ours) | 220.499 | 0.7559 | 0.0738 |
Our modified C-GAN reduces FID by 36.2% and improves SSIM by 15.5% over Neural Style Transfer, producing high-quality, high-resolution colorizations visually close to human-drawn artwork.
Each triplet: input sketch (left), ground truth (middle), generated output (right).
The best-performing model uses a U-Net generator with an N=70 PatchGAN discriminator:
Generator (U-Net): 8-block encoder-decoder with skip connections.
- Encoder:
Conv2D(k=4, s=2)→BatchNorm→LeakyReLU(0.2)per block. Filter progression: 64 → 128 → 256 → 512 → 512 → 512 → 512 → 512, producing a 1×1 bottleneck. - Decoder:
Conv2DTranspose(k=4, s=2)→BatchNorm→ReLUwith skip connections concatenating encoder features. Dropout(0.5) on the first 3 decoder blocks for regularization. - Output:
Conv2DTranspose(3, tanh)producing 256×256×3 RGB images in [-1, 1].
Discriminator (PatchGAN): Classifies overlapping 70×70 patches rather than the full image, enforcing high-frequency structural correctness.
- Input: 6-channel concatenation of
[sketch, target/generated]. - Architecture: 3 downsampling blocks (64 → 128 → 256) →
Conv2D(512, s=1)→BatchNorm→Conv2D(1)outputting a 30×30 patch-level classification map.
Composite Loss:
L(G, D) = L_cGAN(G, D) + lambda_L1 * L_L1(G) + lambda_tv * L_tv(G)
| Loss Term | Formulation | Weight |
|---|---|---|
| Adversarial | E[log D(x,y)] + E[log(1 - D(x, G(x,z)))] |
1.0 |
| L1 Reconstruction | E[‖y - G(x,z)‖₁] |
100.0 |
| Total Variation | sum(|y_{i+1,j} - y_{i,j}| + |y_{i,j+1} - y_{i,j}|) |
1e-4 |
Enables unpaired sketch-to-color translation with two generator-discriminator pairs:
- Generators G, F: U-Net with Instance Normalization (replacing BatchNorm for style-invariant feature normalization). G: sketch → color, F: color → sketch.
- Discriminators D_X, D_Y: PatchGAN with Instance Normalization.
Loss: Adversarial + cycle consistency (lambda_cyc = 10) + identity loss (0.5 * lambda_cyc):
L(G, F, D_X, D_Y) = L_GAN(G, D_Y) + L_GAN(F, D_X) + lambda_cyc * L_cyc(G, F)
L_cyc(G, F) = E[‖F(G(x)) - x‖₁] + E[‖G(F(y)) - y‖₁]
- Optimization-based NST: Iteratively optimizes pixel values of the stylized image using a frozen VGG19 backbone. Content features from
block5_conv2; style features via Gram matrices fromblock{1..5}_conv1. - Fast NST: Single forward pass through Google Magenta's Arbitrary Image Stylization v1-256 network via TensorFlow Hub.
Anime Sketch Colorization Pair (Kaggle)
| Split | Images | Resolution | Format |
|---|---|---|---|
| Train | 14,224 | 256 × 256 | Paired sketch-color PNG |
| Test | 3,545 | 256 × 256 | Paired sketch-color PNG |
| Total | 17,769 |
Preprocessing pipeline:
- Original 512×1024 images split at midpoint into sketch and color channels
- Both resized to 256×256 (nearest-neighbor interpolation)
- Pixel values normalized to [-1, 1]:
(pixel / 127.5) - 1 - Augmentation: Resize to 286×286 → random crop to 256×256 + random horizontal flip
- Shuffled per epoch with buffer size 400
| Hyperparameter | Pix2Pix | CycleGAN | NST |
|---|---|---|---|
| Optimizer | Adam | Adam | Adam |
| Learning rate | 2e-4 | 2e-4 | 0.02 |
| beta_1, beta_2 | 0.5, 0.999 | 0.5, 0.999 | 0.99, 0.999 |
| epsilon | 1e-7 | 1e-7 | 0.1 |
| Batch size | 32 | 8 | 1 |
| Epochs | 150 | 150 | 1,000 |
| Checkpoint freq | 5 epochs | 5 epochs | -- |
| Infrastructure | AWS EC2 + Colab GPU | AWS EC2 + Colab GPU | Colab GPU |
GANime/
├── models/
│ ├── networks.py # U-Net generator & PatchGAN discriminator
│ ├── Pix2PixModel.py # C-GAN training loop with L1 + TV loss
│ ├── CycleGANModel.py # Dual-generator cycle-consistent training
│ ├── NeuralStyleTransferModel.py # VGG19 Gram-matrix optimization
│ └── FastNeuralStyleTransferModel.py # TF Hub single-pass stylization
├── dataloaders/
│ ├── Pix2PixDataLoader.py # Paired image loading & augmentation
│ ├── CycleGANDataLoader.py # Unpaired domain loading & augmentation
│ └── NeuralStyleTransferDataLoader.py
├── utils/
│ ├── evaluation_metrics.py # FID (InceptionV3) & SSIM computation
│ ├── preprocess_data.py # Dataset splitting & format conversion
│ └── download_data.py # Kaggle API dataset fetcher
├── options/
│ ├── TrainOptions.py # Training CLI arguments
│ ├── TestOptions.py # Inference CLI arguments
│ └── EvaluateOptions.py # Evaluation CLI arguments
├── train.py # Unified training entrypoint
├── test.py # Inference & output generation
└── evaluate.py # Metric computation (FID / SSIM)
TensorFlow >= 2.1.0
TensorFlow Hub
NumPy
Matplotlib
SciPy
Kaggle API
# Download via Kaggle API
python utils/download_data.py
# Preprocess for target model
python utils/preprocess_data.py --model pix2pix # paired format
python utils/preprocess_data.py --model cyclegan # unpaired format# Pix2Pix (recommended -- best results)
python train.py --model pix2pix --epochs 150 --lr 2e-4 --batch-size 32 \
--use-tv-loss --lambda-tv-loss 1e-4 \
--data-path <data_dir> --output-path outputs/ --checkpoint-path checkpoints/
# CycleGAN
python train.py --model cyclegan --epochs 150 --lr 2e-4 --batch-size 8 \
--data-path <data_dir> --output-path outputs/ --checkpoint-path checkpoints/
# Neural Style Transfer
python train.py --model neural_style_transfer --epochs 1000 \
--content-path <content_img> --style-path <style_img> --output-path outputs/
# Fast Neural Style Transfer (single-pass, no training required)
python train.py --model fast_neural_style_transfer \
--content-path <content_img> --style-path <style_img> --output-path outputs/Add --resume to continue training from the latest checkpoint.
python test.py --model pix2pix --data-path <data_dir> \
--output-path outputs/ --checkpoint-path checkpoints/# Frechet Inception Distance
python evaluate.py --model pix2pix --metric fid --output-path outputs/
# Structural Similarity Index
python evaluate.py --model pix2pix --metric ssim --output-path outputs/FID uses InceptionV3 (ImageNet-pretrained) activations to measure distributional similarity between generated and real images. SSIM measures per-pixel structural fidelity (luminance, contrast, structure) with filter_size=11, k1=0.01, k2=0.03.
- C-GAN dominates across both metrics due to paired supervision and the L1 reconstruction objective, which provides strong pixel-level gradients that stabilize adversarial training.
- Total variation regularization yields a further 3.3% FID improvement by suppressing high-frequency artifacts and color bleeding at region boundaries.
- CycleGAN produces reasonable colorizations despite unpaired training, but cycle consistency alone is insufficient to match paired supervision quality.
- SSIM plateaus early (~epoch 10) while FID continues improving until ~epoch 35, with qualitative improvements persisting through epoch 100 -- indicating that perceptual quality diverges from pixel-level metrics at later stages of training.
- The PatchGAN discriminator's patch-level classification enforces high-frequency detail accuracy, particularly for hair strands, eye highlights, and clothing folds.
@article{vu2025ganime,
title={GANime: Generating Anime and Manga Character Drawings from Sketches with Deep Learning},
author={Vu, Tai and Yang, Robert},
journal={arXiv preprint arXiv:2508.09207},
year={2025}
}- Gatys, L. A., Ecker, A. S., & Bethge, M. (2015). A Neural Algorithm of Artistic Style. arXiv:1508.06576.
- Ghiasi, G., et al. (2017). Exploring the Structure of a Real-Time, Arbitrary Neural Artistic Stylization Network. arXiv:1705.06830.
- Isola, P., et al. (2017). Image-to-Image Translation with Conditional Adversarial Networks. CVPR 2017.
- Zhu, J.-Y., et al. (2017). Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks. ICCV 2017.
This project is licensed under the MIT License. See LICENSE for details.

