Adding Conditional Control to Text-to-Image Diffusion Models
Introduction
Does previous prompt-based control satisfy our needs?
1.
๋ฐ์ดํฐ์
a.
Specific-domain์์ ์ฌ์ฉ ๊ฐ๋ฅํ ๋ฐ์ดํฐ ๊ท๋ชจ๊ฐ ์ผ๋ฐ์ ์ธ ์ด๋ฏธ์ง-ํ
์คํธ dataset ๋งํผ ํฌ์ง ์์
b.
์ ์ ๋ฐ์ดํฐ์์๋ overfitting ๋ง๊ณ , ์ผ๋ฐ์ ์ธ ์์ฑ ๋ฅ๋ ฅ ์์ง ์๋ ํ์ต ๋ฐฉ๋ฒ ํ์
2.
์ธํ๋ผ
a.
large computation cluster ๋ถ๊ฐ๋ฅํ ๊ฒฝ์ฐ ๋ง์
b.
์ผ๋ฐ์ ์ผ๋ก ๊ฐ์ฉํ ์ ์๋ device(GPU 1๊ฐ ๋ฑ)์์ large model์ ๋น ๋ฅด๊ฒ optimize ํ ์ ์์ด์ผ ํจ
3.
๋ค์ํ input
a.
์ด๋ฏธ์ง ์ฒ๋ฆฌ์๋ pose-to-human, depth-to-image ๋ฑ ๋ค์ํ task๊ฐ ์๊ณ , ์ด๋ฌํ task-specific input condition์ coverํ ์ ์๋ end-to-end learning method ํ์
Method
1. ControlNet
โข
locked copy
โฆ
pre-trained network, locked
โฆ
avoid overfitting, preserve the production-ready quality
โข
trainable copy
โข
zero convolution
โฆ
1x1 convolution layer with both weight and bias initialized as zeros
โฆ
deep feature์ ์๋ก์ด noise ์ถ๊ฐํ์ง ์๊ธฐ ๋๋ฌธ์, ์ถ๊ฐํ layer๋ฅผ scratch(random init) ๋ถํฐ ํ์ตํ ๋ ๋ณด๋ค ๋ ๋น ๋ฅด๊ฒ fine tuning ๊ฐ๋ฅ
โข
zero init ํ๋ฉด update ์ ๋๋๊ฑฐ ์๋๋? ์์ โ ๋
ผ๋ฌธ
2. ControlNet in Image Diffusion Model
diffusion model
stable diffusion
โข
locked original network๋ update ๋์ง ์๊ณ , ์ถ๊ฐํ ControlNet ๋ง update
โฆ
๊ธฐ์กด๋ณด๋ค ์ ์ cost
โ GPU memory, training time ๊ฐ์ โ ์ธํ๋ผ ๋ฌธ์ ํด๊ฒฐ
โฆ
freeze ํ network ๋๋ถ์ ๊ธฐ์กด ์์ฑ ๋ฅ๋ ฅ ์์ง ์๊ณ ์ ์ ๋ฐ์ดํฐ์
๋ง์ผ๋ก fine tuning ๊ฐ๋ฅ
โ ๋ฐ์ดํฐ ๋ฌธ์ ํด๊ฒฐ
โข
Prompt ์ธ์ Condition input ์ฒ๋ฆฌ ๊ฐ๋ฅ โ ๋ค์ํ input ๋ฌธ์ ํด๊ฒฐ
Implementation
โข
default prompt : โa professional, detailed, high-quality imageโ
โข
Automatic prompt : default prompt๋ก ๋ง๋ ์ด๋ฏธ์ง โ BLIP(image captioning model) ์ด์ฉํด prompt ์์ฑ
โข
User prompt : user๊ฐ ๋ง๋ค์ด์ ์ค
1. Canny Edge
2.Hough Line
3. HED Boundary
4.User Sketching
5.Human Pose
6.Semantic Segmentation
Depth
Normal Maps
Cartoon Line Drawing





