Choose how you want to start选择你想以何种方式开始。
Begin with a text prompt when you need a new concept. Upload a reference image when you want to transform or develop an existing visual.
当你需要一个新的概念时,可以从一段文字提示开始。而当你想要对现有视觉元素进行改造或扩展时,则可以上传一张参考图片。

How to Use Grok Image 2.0
Learn how to generate a new image, edit an existing visual, choose practical settings, and refine results with a repeatable creative process.
Begin with a text prompt when you need a new concept. Upload a reference image when you want to transform or develop an existing visual.
当你需要一个新的概念时,可以从一段文字提示开始。而当你想要对现有视觉元素进行改造或扩展时,则可以上传一张参考图片。

Define the subject, setting, composition, style, light, necessary text, and the constraints that matter to your project.
明确你的项目所涉及的主题、场景设定、构图方式、风格特点、光线效果、必要的文字内容,以及那些对作品至关重要的约束条件。

Select an aspect ratio for the final placement and request only as many variations as you need to compare productively.
选择最终布局的纵横比,并根据需要制作尽可能多的不同版本以供对比。

Evaluate the result against the brief. Keep successful decisions and make the next instruction focused on the most important improvement.
将结果与需求规格进行比较评估。保留那些有效的决策,并针对最重要的改进方面给出下一步的指引。

A productive Grok Image 2.0 workflow begins with one decision: are you creating a new visual or changing something that already exists? The answer determines which mode to use, how to structure the prompt, and how you should judge the result.
Text to Image starts with language. It is suitable for concept art, product scenes, campaign directions, illustrations, posters, and any task where no source picture must be preserved. Image Editing starts with both an uploaded picture and an instruction. It is appropriate when a product, person, object, composition, or visual identity already exists and the task is to change part of it.
Before submitting a request, decide where the image will appear. A wide website banner requires different composition from a vertical story. Decide which element must attract attention, how much empty space is needed, and whether any wording must appear in the image. These choices are more useful than adding a long list of decorative adjectives.
Choose Text to Image when the main source material is an idea. Begin by writing the intended subject and its purpose. “A coffee maker” identifies an object but gives the system little direction. “A premium countertop coffee maker photographed for a minimalist ecommerce launch” establishes the subject, commercial purpose, and general creative direction.
Add the environment only when it helps the image. A clean studio background communicates a different message from a lived-in morning kitchen. Then describe framing, light, and style. If the visual needs copy, quote the exact wording and explain where it belongs. If it must leave space for website text added later, request negative space rather than asking the model to invent more content.
Choose Image Editing when the uploaded picture contains something valuable that should guide the next result. The editing instruction should separate changes from invariants. For example, “Place this backpack in a bright airport terminal; keep its shape, logo, materials, zipper placement, and color unchanged” gives the system both an action and boundaries.
Use a clean source whenever possible. Heavy compression, tiny subjects, clipped edges, complex reflections, and existing artifacts reduce the useful information available to an editor. Confirm that you have permission to upload and modify the image, especially when it contains a recognizable person, protected design, private information, or third-party brand.
Replace the plain wall with a softly lit modern studio. Keep the chair, its proportions, green upholstery, wooden legs, camera angle, and foreground shadow unchanged.
A useful prompt acts like a short creative brief. It does not need to describe every pixel. It should identify the decisions that would otherwise be guessed: subject, environment, composition, style, lighting, exact text, and constraints. Put the most important ideas first and make their relationship clear.
Name the primary subject precisely. “Shoes” can mean a pair, a single shoe, footwear on a model, or an overhead collection. “A pair of white trail-running shoes, shown as the hero product for an outdoor campaign” removes that ambiguity. Purpose influences polish, framing, and visual hierarchy, so mention whether the output is a product listing, poster, editorial image, story illustration, or presentation cover.
Describe where the subject appears and how the frame is organized. Useful composition language includes close-up, eye-level view, top-down arrangement, wide establishing shot, centered product, rule-of-thirds placement, symmetrical layout, and negative space on a named side. Avoid incompatible directions such as demanding both an extreme close-up and a wide environment unless you explain a multi-panel layout.
Choose one coherent style direction, then describe the light that supports it. A realistic commercial photograph with soft window light is clearer than a list combining photorealistic, watercolor, anime, and pencil sketch. For text inside an image, put the exact copy in quotation marks, define the headline and supporting hierarchy, and keep the amount manageable. Always proofread generated lettering before publication.
Constraints say what must remain true. In generation they can specify “one product only,” “no people,” or “leave the upper-left corner uncluttered.” In editing they should protect identity, product geometry, branding, pose, camera position, or any detail that already works. For a deeper treatment and complete examples, use the Grok Image 2.0 Prompt Guide.
A premium editorial photograph of a brushed-steel desk lamp on a walnut table, centered three-quarter view, warm evening interior, soft directional light from the left, realistic materials, restrained shadows, generous empty space on the right for website copy, 16:9 composition, no people and no additional products.
Aspect ratio defines the shape of the canvas rather than its final file size. Choosing it before generation helps the model organize the subject for the intended placement. Cropping a square design into a vertical story later may cut off important details, while starting with the correct ratio makes space part of the original composition.
1:1 works for square social posts, profile-style graphics, thumbnails, and balanced product images. 4:3 provides a familiar landscape frame for general photography, presentations, blog graphics, and scenes that need more context without becoming cinematic.
9:16 is designed for full-screen mobile stories, reels, and vertical advertising. 3:4 is useful for portraits, posters, editorial covers, and ecommerce listings that need height but not an extremely narrow frame.
16:9 suits video thumbnails, presentation covers, website sections, and cinematic scenes. 2:1 is a wider banner format that benefits from intentional subject placement and generous negative space.
Consider text overlays, safe areas, mobile crops, and responsive placement. If the same concept must serve several channels, generate or edit for each important ratio instead of assuming one aggressive crop will work everywhere.
Generation is the start of evaluation, not the end of the process. Compare the output with the actual brief rather than asking only whether it looks attractive. A beautiful image can still fail if the product is wrong, the text is misspelled, the intended focal point is weak, or the composition cannot accommodate the final layout.
When requesting multiple outputs, compare them with the same criteria. Do not select a result only because it is the most dramatic. The strongest option is the one that satisfies the useful constraints while leaving the fewest problems to repair.
Upload a JPG, PNG, or WebP source within the limit shown by the interface. Then describe the intended change in outcome-oriented language. Instead of explaining a sequence of software actions, state what the final image should show. “The bottle is on a pale limestone pedestal in a sunlit studio” is usually clearer than a long list of masking operations.
Follow the change with preservation instructions. Identify product shape, labels, colors, facial identity, pose, camera angle, foreground objects, or any other element that should not drift. If a change affects only one region, name it. A focused request is easier to judge and repeat.
Trying to replace a background, change clothing, add text, alter lighting, resize the canvas, and switch artistic style in one request creates competing priorities. Start with the biggest structural change. Save the acceptable output, then use a second pass for a smaller improvement. Versioning makes it possible to return to a good stage if a later edit damages something important.
Do not assume the editor knows which details you value. Repeat the important invariants in each relevant pass. If a product label is correct, explicitly preserve it. If identity matters, state that facial structure, expression, skin tone, hair, and pose should stay consistent. Even then, inspect the result rather than treating preservation as guaranteed.
Effective refinement converts an observation into one clear instruction. “Make it better” does not explain the problem. “Reduce the harsh shadow behind the product and use softer daylight while preserving the composition” identifies both the defect and the allowed change.
Ask for softer contrast, a warmer color temperature, a stronger rim light, reduced reflections, or a more consistent shadow direction. Preserve mood and composition when they already work.
Name the new environment, desired depth, and level of detail. Tell the editor whether the foreground contact shadow and camera perspective should remain.
Request more breathing room, a lower camera angle, a tighter crop, or movement toward a specific third of the frame. Connect the request to the final format.
Reduce the copy, quote exact wording, define headline priority, and reserve clean space. If exact lettering remains critical, plan to add final typography in a design tool.
A vague prompt leaves the most important choices unresolved. Add the intended use, subject detail, frame, and light before adding decorative language. Specificity should reduce ambiguity, not merely increase word count.
Several unrelated styles can produce an inconsistent surface, lighting logic, or composition. Select one primary medium and one supporting influence. If you want alternatives, generate separate directions instead of forcing them into a single frame.
Editing prompts often explain only what should change. Without boundaries, features that already work can drift. State the critical invariants every time they matter.
Dense paragraphs inside generated art are difficult to control and proofread. Prioritize a short headline and essential supporting information. Use external typesetting for legal copy, prices, dates, and other characters that must be exact.
A composition created for one shape may not survive another. Choose the publishing format before generation, and mention placement needs such as space for navigation, product copy, or a mobile safe area.
The provider connection, authentication, credit system, or safety controls may not yet be enabled. Do not repeatedly submit requests. Review the visible status and return when configuration is available.
Confirm the file is JPG, PNG, or WebP and remains under the displayed size limit. Export the source again if its extension does not match its actual format. Avoid damaged files and unusual embedded profiles.
Provider demand can affect completion time. Wait before resubmitting the same prompt. A timed-out browser poll does not always prove that upstream processing stopped, so uncontrolled retries can create duplicate tasks.
Shorten the request to its essential hierarchy. Move the subject and composition earlier, remove conflicting styles, and test one correction. If wording inside the image is wrong, reduce the amount of text and verify every new result.
When account and credit features are enabled, balance changes should be performed on the server and recorded in an auditable ledger. Provider-confirmed failures should follow the Refund Policy. Final package pricing and availability remain visible on the Pricing page.
Practical answers about generating, editing, choosing formats, and resolving common problems.
Choose Text to Image, write a prompt that defines the subject and the most important visual decisions, select an aspect ratio and image count, then submit the request. Review the first result for subject accuracy, composition, lighting, style, and unwanted details before refining the instructions.
Choose Image Editing, upload a supported source image, and describe both the change you want and the elements that must remain consistent. A focused instruction such as changing one background while preserving the product is generally easier to evaluate than several unrelated edits in one request.
The current interface accepts JPG, PNG, and WebP files up to the size shown beside the upload control. A clear, high-quality source with an obvious subject usually provides better editing context than a heavily compressed, blurred, or extremely small image.
Use 1:1 for square posts and thumbnails, 16:9 for widescreen presentations and video covers, 9:16 for stories and mobile creative, 4:3 for general landscape imagery, 3:4 for portraits and product listings, and 2:1 for wide banners.
Completion time depends on provider demand, the selected workflow, and the number of requested images. The interface reports the task state while it runs. If a request exceeds the polling window, wait before submitting the same job again so that duplicate work is not created.
The prompt may contain too many equally important requests, conflicting styles, or an unclear hierarchy. Keep the essential subject and composition near the beginning, remove decorative details that do not matter, and refine one major issue at a time.
Iterative editing is a useful workflow when the connected endpoint supports it. Save each acceptable result, use the strongest version as the next source, and state what must remain unchanged so that successful details are less likely to drift.
Read the displayed error, check the prompt and upload, then try again only after correcting the likely cause. When account and credit features are enabled, provider-confirmed failed or cancelled jobs should be handled by the credit ledger according to the published refund policy.
Open the generator, choose a mode, and build the image through focused iterations.