GPT-6 Astra Writing Review: English, Chinese & Skills | SeaBell
GPT-6 Astra Writing Review: English, Chinese, and Writing Skills
In this small internal writing test, AI evaluators selected GPT-6 Astra more often than GPT-5.6 Sol. The evaluators came from the two model families under test, so family-specific stylistic preferences may have influenced their judgments. Separately, the results showed no consistent benefit from adding a writing skill.
We tested everyday emails, factual explainers, and a short fiction scene in English and Chinese, plus one Spanish email scenario as an exploratory check. The largest differences concerned factual discipline and scene construction. Some drafts treated a plan as an established result or added promises the brief did not authorize; the fiction samples differed in how they showed a character loosening control.
This article draws on an internal set of simulated materials, not real client correspondence or user stories. The full sample set is not yet available for public download; only the excerpts are published here, with brief summaries of relevant context from the complete texts.
🧪 What We Tested
The test produced 56 final texts from 2 models × 2 workflows × 2 task starts × 7 cases. We also retained the corresponding 56 drafts; the pairwise evaluation below concerns the final texts. The seven language-specific case prompts covered Chinese and English versions of three simulated scenarios: a delayed-delivery email, a short team update based on supplied notes, and a scene about siblings closing a repair shop. Spanish reused the email scenario. The test therefore rests on three underlying simulated situations; the 56 texts are not 56 independent tasks.
All included runs took place in Codex Desktop at the high setting. One earlier Astra pilot at the xhigh setting was excluded before evaluation because its configuration did not match the study. The final comparison covers GPT-6 Astra and GPT-5.6 Sol in this setup.
Each model used two workflows:
| Workflow | What happened |
|---|---|
| Without an additional writing skill | The model received the complete task and made an ordinary editorial revision. |
| With an additional writing skill | Chinese drafts used human-writing v1.1.0 for generation and revision. English and Spanish drafts used humanizer v2.8.2 as an editing step. |
The first workflow was an ordinary prompted revision, not a raw output. Both workflows used prompts and an editing pass, and both ran inside Codex Desktop. Their token budgets and editing effort were not matched. In this article, “writing skill” refers to an added set of writing and editing instructions.
All comparisons use matched pairs. For the core Chinese and English cases, each evaluator judged the same 24 Astra-versus-Sol pairs and the same 24 workflow pairs. The model comparison includes 12 pairs from each workflow; the workflow comparison includes 12 pairs from each model. Model and workflow labels were hidden, and A/B order was swapped. The counts describe preferences across those pairs; they are not quality percentages.
| Core comparison, 24 matched pairs | Astra evaluator | Sol evaluator |
|---|---|---|
| Astra vs. Sol, holding workflow constant | Astra 18 / Sol 0 / Tie 6 | Astra 14 / Sol 7 / Tie 3 |
| Added skill vs. no added skill, holding model constant | Skill 3 / No skill 6 / Tie 15 | Skill 11 / No skill 10 / Tie 3 |
Astra's evaluator preferred Astra in 18 pairs and tied six. Sol's evaluator preferred Astra in 14 pairs, preferred Sol in seven, and tied three. The evaluators agreed on 14 of the 24 model pairs: 13 Astra wins and one tie. They differed on the remaining ten.
Evaluators agreed less often in the workflow comparison. They agreed on eight pairs: three ties, four choices for no added skill, and one choice for the added skill. On the other 16 pairs, one evaluator picked a winner while the other tied, or they picked opposite winners. Overall, Astra received more preferences in the model comparison, while the workflow judgments remained mixed.
✉️ Everyday Emails: Tone Is Not the Only Variable
The email prompt was deliberately ordinary. On September 10, a team found inconsistent unit labels in six charts. The written report could still go out on September 12, the charts needed verification by September 15, and the recipient had an internal presentation on September 13. The email had to own the error, tell the recipient not to use unverified charts, ask for a decision by noon on September 11, and offer a 15-minute call if a split-delivery arrangement (the report on the 12th and verified charts by the 15th) would not work.
Only the apologies are quoted below; they differ in tone:
Without an added skill: “I'm sorry for the disruption to your preparation for the September 13 internal presentation.”
With an added skill: “I'm sorry to put this problem in your way as you prepare for your September 13 internal presentation.”
The evaluators tied this pair. The second apology is more personal, but “put this problem in your way” is not a natural expression here. The first is more conventional.
In the second task start, both evaluators preferred Sol's Chinese email with the added skill. Its opening directly acknowledges responsibility; the draft then states the confirmation deadline clearly and offers a fallback call after presenting the recipient's options. It still needs another pass because the delivery dates are repeated unnecessarily.
Before revising tone, confirm that the email states the right dates, includes the warning against using unverified charts, names the decision the recipient must make, and does not overpromise. Tone is secondary to those requirements.
🔎 English and Chinese Explainers: Smooth Prose Can Change a Claim
The explanatory outputs contained the clearest factual overreach. According to the simulated source notes, the team added “Who needs to answer, and by when?” to the template on Wednesday. On Thursday, after two colleagues found long paragraphs hard to scan, the team limited each item to two sentences. The notes never reported that shorter entries had made the question, or anything else, easier to find.
Sol's two English versions from the first task start handled that distinction differently:
Without an added skill: “That should make conditions and requests easier to find.”
With an added skill: “Shorter entries made the new ownership prompt easier to find.”
The sentences change both the object and the certainty of the claim: the first treats easier retrieval as an expectation, while the second reports an unmeasured improvement as fact. The first is more cautious; the second overreaches. Both evaluators preferred the first. Because the drafts began differently and both were revised, no single instruction in a skill file can be isolated as the cause.
A related problem appears in the Chinese version of the scenario. The added-skill draft said:
它到当天仍未解决,可相关同事已经知道问题由谁接手,也知道不能把沉默当成许可。
Editorial English gloss: “As of that day, the issue remained unresolved, but the relevant colleagues already knew who was responsible for it and understood that silence could not be treated as permission.” The Chinese opening “它到当天仍未解决” is clumsy; something closer to “截至当天,这个问题仍未解决” would read more naturally.
More important than the wording is the factual problem: the source notes established only that the question had been passed to procurement and remained unresolved. They did not establish that the relevant colleagues knew who was handling it or that silence would not count as approval. Both evaluators preferred the no-added-skill version. Because the English and Chinese examples reuse the same scenario, they reflect a single underlying pattern.
In the second English start under the added-skill workflow, Astra wrote, “The limit is in place; whether shorter entries give readers enough context still needs checking.” Sol wrote, “The written format gave us a useful record, and the team made it easier to scan.” Both evaluators preferred Astra. Astra leaves the question open; Sol reports a benefit nobody measured.

📻 Fiction: The Radio, the Request, and What Gets Left Out
For the fiction task, the prompt asked for a 500-to-650-word scene about Shen Qing and her younger brother, Shen He, packing their late father's radio in a closing repair shop. Qing is leaving for another city after three months apart. The scene had to show Qing becoming less controlling through the way the siblings handled the radio.
In Astra's second English task start, the two workflow versions show that shift in different ways. In the no-added-skill version, Qing stops herself from directing her brother and makes space for him:
“She had already opened her mouth to tell him which one. Instead she moved the bag a few inches.”
In the added-skill version, the change comes through dialogue after a command:
“Wrap it,” she said. Then, before he could reach for the newspaper, she added, “Will you?”
The loose knob and the paper beneath the radio force them to coordinate. Against that setup, Qing moves the bag instead of finishing the instruction she had begun, giving him room to act. The other version shows the same change through a softened request. Astra's evaluator marked the pair as a tie, while Sol's evaluator preferred the added-skill version.
Astra's Chinese workflow pair drew another split decision. The no-added-skill scene has Qing stop herself from dictating an arrival time:
她差点说七点必须到,停了停。“你几点能来?”
Editorial English gloss: “She almost said that he had to arrive at seven, paused, and asked, ‘What time can you come?’”
The added-skill scene ends with Qing giving one practical instruction and then stepping aside:
“横着。”她说,随后往旁边让了让。
Editorial English gloss: “‘On its side,’ she said, then stepped aside.”
Astra's evaluator marked a tie; Sol's evaluator preferred the no-added-skill scene. Earlier, Shen He had shown Qing how to steady the loose knob. When Qing steps aside, she gives him room to finish packing the radio rather than simply following her instruction.
Sol's first English scene without the extra skill is built around small physical concessions. Qing directs her brother to use both hands, then catches the other side of a heavy box. Later, when he suggests the radio could help them stay connected, she says, “You have a phone.” Shen He replies, “That's not what I said.” The scene ends with Qing putting the tape beside him while keeping the cutter in her own hand. Astra's evaluator tied this model pair; Sol's evaluator selected Sol.

Sol produced eight English and Chinese final fiction texts (two workflows × two task starts × two languages); three omitted a required background detail: the father had died the previous year. Omitting that detail fails a prompt requirement without necessarily undermining the scenes themselves.
🧩 What the Core Cases Suggest About Writing Skills
Of the 12 workflow pairs where Astra was the underlying model, Astra's evaluator marked every pair as a tie; Sol's evaluator chose the skill seven times, no added skill three times, and tied twice. The other 12 pairs, where Sol was the underlying model, were more mixed: Astra's evaluator chose the skill three times, no added skill six times, and tied three times; Sol's evaluator chose the skill four times, no added skill seven times, and tied once.
The case-level results explain the split. Sol's Chinese email with the skill was more direct, while its English explainer turned a predicted benefit into a reported result. In fiction, both workflows offered plausible ways to show Qing becoming less controlling.
A Brief Spanish Check
The Spanish comparison covered only the email scenario, which provided too little material for a language-level ranking. Of the four model pairs, Astra's evaluator chose Astra twice and tied twice; Sol's evaluator chose Astra three times and Sol once. For the four workflow pairs, Astra's evaluator chose no added skill once and tied the other three. Sol's evaluator chose the skill once, no added skill twice, and tied once.
The dates and the warning against using unverified charts still needed checking in the Spanish drafts. Pronoun and register choices gave the evaluators room to disagree. No Spanish native-speaker editor reviewed the samples, and the test included no Spanish-specific editing workflow.
🛠️ How to Use This Result on Your Own Draft
Begin with complete facts and a clear audience, then produce a normal revised draft. If an additional writing skill suits the task, run a second version and compare both drafts against the source material. For correspondence, verify every commitment as though you must carry it out. For an explainer, separate observed results from predictions. For fiction, read the full scene before judging an isolated polished sentence.
📚 Long-Form Fiction Is Outside This Test
This comparison covers one short scene, not chapter-to-chapter continuity. Longer projects need a way to keep character notes, world details, and earlier decisions available during drafting. SeaBell is one such workspace; this test did not assess whether it offers Astra or these specific skills.
For related reading, see the Sol and Fable fiction comparison and our chapter-based AI writing workflow.
📋 Method and Limits
This internal workflow comparison used anonymized pairs with A/B order swapped. Each evaluator received the supplied assignments, source notes, complete sample packet, pair definitions, and rubric. Each requested comparison presented two texts as A and B; the evaluator chose between the two texts or declared a tie, and the mapping from A/B back to model and workflow was applied only after judging. The evaluators were not told which model produced which text. Because the evaluators came from the two model families under test, their choices could reflect family-specific stylistic preferences or different thresholds for ties. The article itself was reviewed with AI assistance, and its comments on Chinese and English phrasing are editorial readings rather than findings from a native-speaker panel. These results come from one small internal sample, not a statistically representative benchmark.
For each model-workflow combination, the seven cases were run together in each of two task starts, so cases within a batch shared context. The six Chinese and English core cases reuse three simulated scenarios; Spanish reuses the email scenario. Earlier skill instructions could remain in a batch context, and reversing the case order in the second task start did not eliminate that possibility. Two Sol English explainers were length-corrected after their initial drafts were saved. The final texts met their stated length ranges and included those revisions.
The skill workflows were not equivalent across languages. Chinese used human-writing v1.1.0 for writing and revision; English and Spanish used humanizer v2.8.2 for editing. The Chinese-focused skills shuorenhua and humanizer-zh were outside this comparison. These differences prevent the test from isolating the effect of any one file or establishing that the editing budgets were matched. Cost, speed, AI-detector results, and publication readiness were not measured.
The side-by-side excerpts illustrate selected findings from the larger internal record. They show where a revision preserved the source, altered a claim, or drew disagreement between evaluators. Because the full prompts, source notes, complete drafts, and pairwise inputs are not public, readers cannot independently verify every factual or scene-level judgment discussed here. Any broader ranking would require more independent tasks and human evaluation.
Related Posts

GPT-5.6 Sol vs Claude Fable 5 for Creative Writing | SeaBell
We ran GPT-5.6 Sol and Claude Fable 5 through six identical fiction prompts for voice, dialogue, continuity, revision, and scene pressure.

Gemini 3.5 Flash for creative writing: a practical novel workflow
A practical guide to Gemini 3.5 Flash for creative writing, with a novel workflow for premise testing, character consistency, chapter drafting, and revision.
