How text-to-image generation actually works

Modern AI image generators are built on diffusion models: the system is trained to remove noise from images, learning from pairs of images and their text descriptions. At generation time, it starts from pure noise and repeatedly "denoises" toward a picture that matches your prompt — the technique behind the latent diffusion models that made Stable Diffusion famous [1]. That's why the same prompt produces different results each run, and why prompt wording ("cinematic lighting", "85mm lens") changes output far more than intuition suggests. OpenAI's line illustrates the evolution: DALL·E (January 2021), DALL·E 2 (April 2022) [2], and DALL·E 3 (September 20, 2023) [3].

The five generators that matter, compared

Midjourney — the aesthetic favorite

Founded by David Holz, Midjourney iterates fast: V5 (Mar 2023), V6 (Dec 2023), V7 (April 3, 2025 — its first new base model in nearly a year) [4], and the V8 alpha (March 17, 2026), with the V8 edit model shipping August 27, 2026 [5]. Pricing as of May 2026: Basic $10 (≈200 generations), Standard $30, Pro $60, Mega $120 [6][7]. Two structural facts shape every "best ai image generator" verdict: Midjourney has no public API — everything runs through Discord or the web app [7] — and V7 launched without full feature parity (image editing arrived later) [4].

DALL·E / GPT Image (OpenAI) — the integrated option

After DALL·E 3, OpenAI folded image generation into ChatGPT itself: ChatGPT Image on GPT-4o launched March 25, 2025 — and demand was so high OpenAI's status page joked "our GPUs are melting" while rate-limiting requests [8] — followed by the gpt-image-1 API (April 23, 2025) [9] and gpt-image-2 in the API and Codex (April 21, 2026) [10]. Consumer access rides the ChatGPT plans (Free, Go $8, Plus $20, Pro $200) [11]; API pricing runs roughly $0.02–$0.19 per image by quality tier [12]. If you already pay for ChatGPT, this is the "ai image generator free" answer — the free tier includes image generation with daily limits.

Stable Diffusion (Stability AI) — the open-weights workhorse

The timeline: SD 1.4 (Aug 2022), SDXL (Jul 2023), SD 3 (June 12, 2024 — pulled after quality backlash, memorably documented by Ars Technica's "body horror" headline) [13], and SD 3.5 (October 22, 2024) [14]. Weights are free under Stability's community license (with commercial revenue thresholds); the hosted API has a 25-credit free tier [15]. Reports of a "Stable Diffusion 4" remain unverified as of September 2026 [16]. It's the choice when you want to run generation on your own hardware — see our local AI guide for the setup.

Flux (Black Forest Labs) — the quality challenger

Founded in August 2024 by ex-Stability researchers, BFL shipped FLUX.1 schnell/dev/pro (Aug 2024) [17], FLUX.2 (December 1, 2025) [18], and FLUX 3 Video with native audio (August 4, 2026) [19]. FLUX.2 [dev] is non-commercial but costs only ≈$0.012/image via API hosts [18]; FLUX 3 Video runs $0.06/s draft up to $0.29/s FHD [19]. The $300M Series B at a $3.25B valuation (Dec 2025) says the market believes in it [20].

Ideogram — text that actually renders

Toronto's Ideogram (ex-Google Brain founders) has been the reference for rendering readable text inside images — signs, posters, logos — since 1.0 (Aug 2023), with 2.0 (Aug 2024) and 3.0 (Mar 2025) built around realism and design fidelity [21]. Pricing as of May 2026: Free (≈25 generations/day), Basic $8, Plus $20 [22]. If your prompts involve words on the image, try Ideogram before anything else.

AI passport photos: the rules most people get wrong

The "how to use ai to make a passport photo" workflow is straightforward — pose against a plain background, use an AI tool to crop, center and replace the background — but the compliance part is where applications fail. Official rules in many countries prohibit digitally altered photos. The US State Department, for example, requires passport photos that have not been digitally altered or retouched, including background removal that changes your appearance [23]. Before submitting anything: check your government's photo requirements, keep the original file, and treat AI-generated or heavily edited photos as a rejection risk for official documents. For unofficial ID-style photos (badges, gym cards, social profiles), the ai passport image generator tools work fine — just don't cross the official-document line.

Photo-editing prompts that work

"Best chat gpt prompt for photo editing" and "best gemini ai photo editing prompts" follow the same pattern as our chatbot prompt rules: state the image, the change, and the constraint. "Remove the person in the background, keep lighting natural, don't change my face" beats "make this photo better." ChatGPT's image tools handle inpainting-style edits natively, and Gemini's photo prompts work best when you describe the target style rather than the editing operation.

AI video: from Sora to Veo

The "best ai video generator free" answer in 2026: Sora's free ChatGPT tier or Kling's 66 daily credits, both watermarked. Pay when you need audio polish, higher resolution or commercial rights.

Getty v. Stability AI (UK High Court, Nov 4, 2025): Getty's primary copyright claims were largely abandoned at trial, the secondary copyright claim over model weights was dismissed, and trademark findings were "extremely limited" — a landmark first ruling on AI training in the UK [39][40]. A parallel US case in Delaware continues [40]. Separately, Getty licensed its library to OpenAI in June 2026 [41]. On provenance: OpenAI added C2PA metadata to DALL·E 3 images in February 2024 and Google's SynthID watermarks to ChatGPT images in March 2025 [42]; Google applied C2PA 2.1 credentials in the Gemini app [43]. On the law: the EU AI Act's transparency obligations (machine-readable marking of AI outputs, visible deepfake disclosure) take effect August 2, 2026 [44]; California's AI Transparency Act (SB 942, amended by AB 853) is operative the same day [44]; Tennessee's ELVIS Act has protected voice/likeness since July 1, 2024 [45]; and the US TAKE IT DOWN Act requires platforms to remove non-consensual intimate imagery within 48 hours [46]. Practical rule: generated images carry metadata; screenshots strip it — and passing off generated media as real is exactly what the new transparency laws target.

Comparison table (verified 2026)

ToolFree tierEntry paidBest forWatch out
MidjourneyNo (trial)Basic $10/moAesthetic qualityNo public API; Discord/web only
GPT Image (DALL·E)Yes (daily limits)ChatGPT Go $8/moIntegrated editing in ChatGPTRate limits on lower tiers
Stable DiffusionOpen weightsAPI credits (25 free/mo)Self-hosting, fine-tuningCommunity license limits
Flux (BFL)dev weights (non-commercial)≈$0.012/imageQuality per dollar[pro] weights closed
IdeogramYes, ~25/dayBasic $8/moText inside imagesDaily free ceiling
SoraYes (limited)Go $8/moVideo with audio (Sora 2)Credit caps; EU rollout history
RunwayLimitedfrom $12/moPro video editingCredit-metered; plan restructuring 2026
KlingYes, 66 credits/dayStandard $8.80/moVolume + multi-shotCredits expire monthly
Veo 3Preview via VertexAPI $0.15–0.40/secProduction API videoPer-second billing adds up

Prices verified from vendor pages and reputable trackers, 2026. Sources: [6][11][12][15][18][19][22][27][29][32][35][38].

Frequently asked questions

What is the best free AI image generator in 2026?+
For most people: ChatGPT's image generation on the free tier (daily limits apply) or Ideogram's free tier (~25 generations/day). Stable Diffusion and FLUX [dev] are free if you can run them locally — see our local AI guide.
Is there a free AI passport photo generator?+
Several tools offer free AI passport-photo framing, but official rules usually prohibit digitally altered photos. US passport photos, for example, must not be digitally altered or retouched [23]. Always check your government's photo rules before submitting.
Are AI-generated images copyrighted?+
It depends on jurisdiction and human authorship, and the law is still developing. In the UK, Getty's copyright claims against Stability AI were largely dismissed in November 2025 [39][40]; a parallel US case continues. Check each tool's terms and your local law before commercial use.
What is C2PA and how do you detect AI images?+
C2PA is an open standard embedding provenance metadata (which tool created an image, and when) in the file. OpenAI adds C2PA metadata and Google SynthID watermarks to generated images [42][43]. Detection tools read this metadata; screenshots strip it.
How much does Midjourney cost per month?+
As of May 2026: Basic $10 (≈200 generations), Standard $30, Pro $60, Mega $120 [6][7]. No public API — generation runs through Discord or the web app.
Is Sora available in Europe?+
Sora 1 skipped the EU and UK at its December 2024 launch [24]. Availability has expanded since, but rollout varies by country — check OpenAI's availability page for your region.

Sources

  1. Rombach et al., "High-Resolution Image Synthesis with Latent Diffusion Models" — arxiv.org/abs/2112.10752
  2. OpenAI, DALL·E 2 — openai.com
  3. Reuters, "OpenAI unveils DALL-E 3", Sep 2023 — reuters.com
  4. TechCrunch, "Midjourney releases its first new AI image model in nearly a year" (V7), Apr 2025 — techcrunch.com
  5. Midjourney updates — updates.midjourney.com
  6. Midjourney docs (pricing) — docs.midjourney.com
  7. Mobbi.ai, Midjourney pricing guide — mobbi.ai
  8. The Verge, "ChatGPT says our GPUs are melting", Mar 2025 — theverge.com
  9. TechCrunch, "OpenAI makes its upgraded image generator available to developers", Apr 2025 — techcrunch.com
  10. OpenAI developer docs, gpt-image-2 — developers.openai.com
  11. OpenAI, Introducing ChatGPT Go — openai.com
  12. OpenAI developer docs, gpt-image-1 pricing — developers.openai.com
  13. Ars Technica, "Ridiculed Stable Diffusion 3 release excels at AI-generated body horror", Jun 2024 — arstechnica.com
  14. Stability AI, "Introducing Stable Diffusion 3.5" — stability.ai
  15. Stability platform pricing — platform.stability.ai
  16. LocalAIMaster, "Stable Diffusion 4" (unverified status) — localaimaster.com
  17. Hugging Face, black-forest-labs/FLUX.1-dev — huggingface.co
  18. LLM Stats, FLUX.2 dev — llm-stats.com
  19. LLM Stats, FLUX 3 Video launch — llm-stats.com
  20. Crunchbase News, Black Forest Labs raise — news.crunchbase.com
  21. Ideogram, Ideogram 3.0 — ideogram.ai
  22. FluxNote, Ideogram pricing guide 2026 — fluxnote.io
  23. US Department of State, passport photos — travel.state.gov
  24. TechCrunch, "OpenAI's Sora video generator might not be available in the EU at launch", Dec 2024 — techcrunch.com
  25. NBC News, "OpenAI announces Sora 2 AI video audio app" — nbcnews.com
  26. Gizmodo, "OpenAI officially launches video generator Sora 2" — gizmodo.com
  27. ZDNET, "How to use OpenAI's Sora" — zdnet.com
  28. ThursdAI, Runway company profile — thursdai.news
  29. Runway pricing — runway.com
  30. Runway Help, "Unlimited plan is switching to Max" — help.runwayml.com
  31. Cliprise, "Pika 2.5 released" — cliprise.app
  32. Pika pricing — pikaslabs.com
  33. Bernama, "Kling 3.0" — bernamabiz.com
  34. Kling release history — kling.ai
  35. Modellix, Kling pricing — modellix.ai
  36. Google blog, "Video & image generation update", Dec 2024 (Veo/Veo 2) — blog.google
  37. Google Cloud, "Veo 3 available for everyone in public preview on Vertex AI" — cloud.google.com
  38. FoneArena, "Google Veo 3 vertical 1080p HD" — fonearena.com
  39. Bird & Bird, "Stability AI defeats Getty Images copyright claims", Nov 2025 — twobirds.com
  40. Cleary Gottlieb, "UK High Court issues landmark ruling in Getty Images v. Stability AI" — clearygottlieb.com
  41. Investing.com (HK), Getty–OpenAI license, Jun 2026 — hk.investing.com
  42. The Next Web, "OpenAI C2PA SynthID AI image detection watermark" — thenextweb.com
  43. Gadgets360, "Google Gemini C2PA OpenAI SynthID watermark" — gadgets360.com
  44. Starling Lab, "Key legislation and regulation on AI and authenticity" — starlinglab.org
  45. Reed Smith, "Tennessee efforts to protect ROP against GenAI imitations/deepfakes" — reedsmith.com
  46. Senator Klobuchar, "TAKE IT DOWN Act signed into law" — klobuchar.senate.gov