Workflow guide
Text to 3D, and why going through an image first gives you more control
Text to 3D in one step is convenient and hard to direct. Generating an image first, then converting it, gives back the control that matters.
Text to 3D is the demo everyone wants: describe an object, receive a mesh. It exists, it works, and in practice most people who use it seriously end up doing something slightly different, because the one-step route removes the only cheap place to iterate.
Mixar deliberately has no one-step text-to-3D button, and the reason is the whole argument of this page: generating an image first, iterating on it cheaply, then converting the one you approved gives you control a single step cannot. (Mixar is a 3D editor built on Blender, with image generation, image-to-3D and an AI agent all inside the one application rather than spread across browser tabs.)
The alternative is two steps. Generate an image from the text, then convert the image to 3D. It sounds like more work and usually is not, for reasons that become obvious the third time you try to nudge a result in a specific direction.
The problem with one step#
When text goes straight to a mesh, every decision happens at once and invisibly. The silhouette, the proportions, the surface detail, the style and the pose are all resolved in a single pass whose only input is your sentence.
That leaves you one lever. If the result is close but the proportions are wrong, you change the words and regenerate, and everything else changes too. If it is right except for the pose, same. Each attempt is a full conversion, which is the slowest and most expensive operation in the chain, and you are paying it to explore a decision that was made in the first few seconds.
There is also a communication limit. Words are a poor medium for shape. "Chunky but not cartoonish", "the proportions of a 1970s desk fan", "wider at the base than that": these are things you can point at instantly and describe badly.
The two-step route#
Generate an image first. Look at it. Iterate on it until the silhouette and proportions are what you want, which is fast and cheap because images are fast and cheap. Then convert the image you chose.
What this buys:
You see the decision before committing to it. The image is the silhouette, the proportions and the style, resolved and visible. Converting a reference you have already approved removes most of the surprise from the expensive step.
Iteration happens at the cheap end. The asymmetry is the whole argument: an image comes back in seconds, a 3D conversion is a queued job measured in minutes and costs several times more. Ten image variations to find the silhouette you want, then one conversion, is a fundamentally different spend from ten conversions to find the same thing. You pay the expensive operation once, on the option you already approved.
You can reference instead of describe. Image generation takes reference images alongside the prompt, so a rough sketch, a photograph of something with the right proportions, or a previous asset from the same project directs the shape far better than an adjective.
You can block out the composition yourself. Depth-guided generation is the strongest version of this: block the shape out with primitives in the viewport, capture the depth, and generate an image that respects your geometry and camera rather than inventing its own. The structure is yours; only the surface is generated. That inverts the usual complaint about generated references being beautiful and structurally useless.
You can supply multiple views. Once you have a reference you like, generating consistent additional angles of it and converting them together moves the hidden surfaces from invention to reconstruction. This is the single largest quality change available, and the one-step route has no equivalent.
One step
Silhouette, proportion, style and pose all resolved at once from one sentence. Adjusting any of them means regenerating all of them, and each attempt pays for a full conversion.
Two steps
Iterate on images, which are fast and cheap, until the silhouette is right. Then spend the expensive operation once, on the reference you approved.
What it costs#
One extra step, and a discipline: the intermediate image has to be a good conversion reference, which is not the same as a good picture.
A good conversion reference is evenly lit with no hard shadows, on a plain contrasting background, at a three-quarter angle slightly above centre, with the whole subject in frame and in focus. A good picture is often dramatically lit, tightly cropped and shallow-focused, all of which make the conversion worse. The same rules apply as when photographing a real object, covered in photo to 3D model.
The other cost is honesty about what you are doing: this is not text to 3D, it is text to image to 3D, and anyone selling you the first thing is selling a simpler story.
How this looks in Mixar#
There is no single text-to-3D button, and the two-step route is why. The pieces are the moodboard and the generation sidebar, in one editor:
Generate images from a prompt
With optional reference images, at 1:1, 16:9, 9:16, 4:3, 3:4 or 21:9, and 1K, 2K or 4K depending on the model. Results land on the moodboard, not in a folder.
Or block out and render
Rough the shape in the viewport with primitives and take the depth-guided path instead. You supply the structure; the model only supplies the surface.
Iterate on the board
Variations sit side by side. Cmd/Ctrl+P sends the one you like straight into the chat as the reference for the next step.
Convert the one you chose
Face count from 40,000 to 1,500,000, PBR textures on or off, and a Geometry type that returns an untextured white model when you intend to author the surface yourself.
Add angles if asymmetric
One frontal image plus up to seven named angles, eight maximum.
The mesh imports into the scene
Where the cleanup pass is the next line of the same brief.
The Geometry option is the one people miss. If the mesh is a blockout whose surface you will author properly, generating a texture you intend to discard costs time and credits for nothing.
Image-to-3D conversion does accept an optional prompt alongside the image, which is useful for disambiguating material or context. It is a hint to a process driven by the picture, not a substitute for one.
When one step is the right call#
Not every case wants the control.
Filling a scene fast. Twenty background objects for a layout test do not need art direction. One step, twenty prompts, move on.
Exploring at the concept stage. When you genuinely do not know what you want yet, the speed of a single step is worth more than the control you are not ready to exercise.
Simple, unambiguous objects. A wooden crate is a wooden crate. There is not much to direct.
The two-step route earns its extra step exactly when you care what the thing looks like, which is also when the one-step route frustrates most.
After the conversion, either way#
The mesh is a detailed blockout regardless of route: dense triangles, machine-generated UVs, arbitrary scale. Clean it, retopologise to a real budget (Blender's free Remesh modifier and QuadriFlow will do it if you have nothing else), rebuild the UVs, bake the detail down, texture over the generated maps, then name and export. That sequence is identical whichever way the mesh arrived and is most of the elapsed time, which is why it is the part worth briefing to an AI agent for Blender rather than doing by hand across a batch.
The full stage-by-stage picture is in AI 3D, stage by stage, and the conversion step specifically is covered in image to 3D.
Frequently asked questions
Does text to 3D work?
It works, in the sense that you will get a mesh that matches your description. What it does not give you is control: every decision about silhouette, proportion, style and pose is resolved in one pass whose only input is your sentence, so adjusting one aspect means regenerating all of them. Most people doing this seriously generate an image first, iterate on that cheaply, then convert the reference they chose.
Is text to 3D or image to 3D better?
Image to 3D, for anything you care about the look of, because you can see and approve the reference before paying for the expensive step. Text to 3D is better when speed matters more than direction: filling a scene for a layout test, early concept exploration, or simple objects where there is nothing to direct. The two-step route costs one extra step and gives back the ability to iterate at the cheap end.
Can I use a sketch instead of a text prompt?
Yes, and it usually works better. Image generation takes reference images alongside the prompt, so a rough sketch or a photograph with the right proportions directs shape far more precisely than adjectives do. Stronger still is depth-guided generation: block the form out with primitives in the viewport, capture the depth, and generate an image that respects your geometry and camera. The structure stays yours.
What makes a good reference image for conversion?
Not the same things that make a good picture. Even, flat lighting with no hard shadows, a plain background with clear contrast against the subject, a three-quarter angle slightly above the object's centre, the whole subject in frame, and everything in focus. Dramatic lighting, tight crops and shallow depth of field all look better and convert worse, which is a genuinely annoying tradeoff the first time you meet it.
How do I go from text to a 3D model in Blender?
The reliable route is two steps rather than one: generate an image from your prompt, iterate on it until the silhouette and proportions are right, then convert the image you approved. In Mixar, a 3D editor built on Blender, both steps live in the same application, so the images land on a moodboard and the converted mesh lands in your scene rather than a downloads folder.