Stable Diffusion 2.0 Release
AI Startup
San Francisco, CAOn-siteFull-time$140k – $220k / year
About this role
Stable Diffusion 2.0 Release — Stability AI
Stable Diffusion 2.0 Release
Nov 24
We are pleased to announce the open-source release of Stable Diffusion Version 2.
The original Stable Diffusion V1 led by CompVis changed the nature of open source AI models and spawned hundreds of other models and innovations worldwide. It had one of the fastest climbs to 10K GitHub stars of any software, rocketing through 33K stars in less than two months.
Source: A16z and Github
The dynamic team of Robin Rombach (Stability AI) and Patrick Esser (Runway ML) from the CompVis Group at LMU Munich, headed by Prof. Dr. Björn Ommer, led the original Stable Diffusion V1 release. They built on their prior work in the lab with Latent Diffusion Models and got critical support from LAION and Eleuther AI. In our earlier blog post, you can read more about the original Stable Diffusion V1 release. Robin is now leading the effort with Katherine Crowson at Stability AI to create the next generation of media models with our broader team.
Stable Diffusion 2.0 delivers several big improvements and features versus the original V1 release, so let’s dive in and take a look at them.
New Text-to-Image Diffusion Models
The Stable Diffusion 2.0 release includes robust text-to-image models trained using a brand new text encoder (OpenCLIP), developed by LAION with support from Stability AI, which greatly improves the quality of the generated images compared to earlier V1 releases. The text-to-image models in this release can generate images with default resolutions of 512x512 pixels and 768x768 pixels.
These models are trained on an aesthetic subset of the LAION-5B dataset created by the DeepFloyd team at Stability AI, which is then further filtered to remove adult content using LAION’s NSFW filter.
Examples of images produced using Stable Diffusion 2.0, at 768x768 image resolution.
Super-resolution Upscaler Diffusion Models
Stable Diffusion 2.0 also includes an Upscaler Diffusion model t