Search

ControlNet

Adding Conditional Control to Text-to-Image Diffusion Models

Introduction

Does previous prompt-based control satisfy our needs?
1.
๋ฐ์ดํ„ฐ์…‹
a.
Specific-domain์—์„œ ์‚ฌ์šฉ ๊ฐ€๋Šฅํ•œ ๋ฐ์ดํ„ฐ ๊ทœ๋ชจ๊ฐ€ ์ผ๋ฐ˜์ ์ธ ์ด๋ฏธ์ง€-ํ…์ŠคํŠธ dataset ๋งŒํผ ํฌ์ง€ ์•Š์Œ
b.
์ ์€ ๋ฐ์ดํ„ฐ์—์„œ๋„ overfitting ๋ง‰๊ณ , ์ผ๋ฐ˜์ ์ธ ์ƒ์„ฑ ๋Šฅ๋ ฅ ์žƒ์ง€ ์•Š๋Š” ํ•™์Šต ๋ฐฉ๋ฒ• ํ•„์š”
2.
์ธํ”„๋ผ
a.
large computation cluster ๋ถˆ๊ฐ€๋Šฅํ•œ ๊ฒฝ์šฐ ๋งŽ์Œ
b.
์ผ๋ฐ˜์ ์œผ๋กœ ๊ฐ€์šฉํ•  ์ˆ˜ ์žˆ๋Š” device(GPU 1๊ฐœ ๋“ฑ)์—์„œ large model์„ ๋น ๋ฅด๊ฒŒ optimize ํ•  ์ˆ˜ ์žˆ์–ด์•ผ ํ•จ
3.
๋‹ค์–‘ํ•œ input
a.
์ด๋ฏธ์ง€ ์ฒ˜๋ฆฌ์—๋Š” pose-to-human, depth-to-image ๋“ฑ ๋‹ค์–‘ํ•œ task๊ฐ€ ์žˆ๊ณ , ์ด๋Ÿฌํ•œ task-specific input condition์„ coverํ•  ์ˆ˜ ์žˆ๋Š” end-to-end learning method ํ•„์š”

Method

1. ControlNet

โ€ข
locked copy
โ—ฆ
pre-trained network, locked
โ—ฆ
avoid overfitting, preserve the production-ready quality
โ€ข
trainable copy
โ€ข
zero convolution
โ—ฆ
1x1 convolution layer with both weight and bias initialized as zeros
โ—ฆ
deep feature์— ์ƒˆ๋กœ์šด noise ์ถ”๊ฐ€ํ•˜์ง€ ์•Š๊ธฐ ๋•Œ๋ฌธ์—, ์ถ”๊ฐ€ํ•œ layer๋ฅผ scratch(random init) ๋ถ€ํ„ฐ ํ•™์Šตํ•  ๋•Œ ๋ณด๋‹ค ๋” ๋น ๋ฅด๊ฒŒ fine tuning ๊ฐ€๋Šฅ
โ€ข
zero init ํ•˜๋ฉด update ์•ˆ ๋˜๋Š”๊ฑฐ ์•„๋‹ˆ๋ƒ? ์ˆ˜์‹ โ†’ ๋…ผ๋ฌธ

2. ControlNet in Image Diffusion Model

diffusion model
stable diffusion
โ€ข
locked original network๋Š” update ๋˜์ง€ ์•Š๊ณ , ์ถ”๊ฐ€ํ•œ ControlNet ๋งŒ update
โ—ฆ
๊ธฐ์กด๋ณด๋‹ค ์ ์€ cost
โ†’ GPU memory, training time ๊ฐ์†Œ โ†’ ์ธํ”„๋ผ ๋ฌธ์ œ ํ•ด๊ฒฐ
โ—ฆ
freeze ํ•œ network ๋•๋ถ„์— ๊ธฐ์กด ์ƒ์„ฑ ๋Šฅ๋ ฅ ์žƒ์ง€ ์•Š๊ณ  ์ ์€ ๋ฐ์ดํ„ฐ์…‹ ๋งŒ์œผ๋กœ fine tuning ๊ฐ€๋Šฅ โ†’ ๋ฐ์ดํ„ฐ ๋ฌธ์ œ ํ•ด๊ฒฐ
โ€ข
Prompt ์™ธ์— Condition input ์ฒ˜๋ฆฌ ๊ฐ€๋Šฅ โ†’ ๋‹ค์–‘ํ•œ input ๋ฌธ์ œ ํ•ด๊ฒฐ

Implementation

โ€ข
default prompt : โ€œa professional, detailed, high-quality imageโ€
โ€ข
Automatic prompt : default prompt๋กœ ๋งŒ๋“  ์ด๋ฏธ์ง€ โ†’ BLIP(image captioning model) ์ด์šฉํ•ด prompt ์ƒ์„ฑ
โ€ข
User prompt : user๊ฐ€ ๋งŒ๋“ค์–ด์„œ ์คŒ
1. Canny Edge
2.Hough Line
3. HED Boundary
4.User Sketching
5.Human Pose
6.Semantic Segmentation
Depth
Normal Maps
Cartoon Line Drawing