SeaArt AI Novel
APP

GPT-6 Astra Writing Review: English, Chinese & Skills | SeaBell

NaronPublished on Sep 8, 2026 8 min read
We tested GPT-6 Astra and GPT-5.6 Sol on English and Chinese emails, explainers, and fiction, with and without writing skills. See the samples and tradeoffs.

GPT-6 Astra Writing Review: English, Chinese, and Writing Skills

In this small internal writing test, AI evaluators selected GPT-6 Astra more often than GPT-5.6 Sol. The evaluators came from the two model families under test, so family-specific stylistic preferences may have influenced their judgments. Separately, the results showed no consistent benefit from adding a writing skill.

We tested everyday emails, factual explainers, and a short fiction scene in English and Chinese, plus one Spanish email scenario as an exploratory check. The largest differences concerned factual discipline and scene construction. Some drafts treated a plan as an established result or added promises the brief did not authorize; the fiction samples differed in how they showed a character loosening control.

This article draws on an internal set of simulated materials, not real client correspondence or user stories. The full sample set is not yet available for public download; only the excerpts are published here, with brief summaries of relevant context from the complete texts.


🧪 What We Tested

The test produced 56 final texts from 2 models × 2 workflows × 2 task starts × 7 cases. We also retained the corresponding 56 drafts; the pairwise evaluation below concerns the final texts. The seven language-specific case prompts covered Chinese and English versions of three simulated scenarios: a delayed-delivery email, a short team update based on supplied notes, and a scene about siblings closing a repair shop. Spanish reused the email scenario. The test therefore rests on three underlying simulated situations; the 56 texts are not 56 independent tasks.

All included runs took place in Codex Desktop at the high setting. One earlier Astra pilot at the xhigh setting was excluded before evaluation because its configuration did not match the study. The final comparison covers GPT-6 Astra and GPT-5.6 Sol in this setup.

Each model used two workflows:

WorkflowWhat happened
Without an additional writing skillThe model received the complete task and made an ordinary editorial revision.
With an additional writing skillChinese drafts used human-writing v1.1.0 for generation and revision. English and Spanish drafts used humanizer v2.8.2 as an editing step.

The first workflow was an ordinary prompted revision, not a raw output. Both workflows used prompts and an editing pass, and both ran inside Codex Desktop. Their token budgets and editing effort were not matched. In this article, “writing skill” refers to an added set of writing and editing instructions.

All comparisons use matched pairs. For the core Chinese and English cases, each evaluator judged the same 24 Astra-versus-Sol pairs and the same 24 workflow pairs. The model comparison includes 12 pairs from each workflow; the workflow comparison includes 12 pairs from each model. Model and workflow labels were hidden, and A/B order was swapped. The counts describe preferences across those pairs; they are not quality percentages.

Core comparison, 24 matched pairsAstra evaluatorSol evaluator
Astra vs. Sol, holding workflow constantAstra 18 / Sol 0 / Tie 6Astra 14 / Sol 7 / Tie 3
Added skill vs. no added skill, holding model constantSkill 3 / No skill 6 / Tie 15Skill 11 / No skill 10 / Tie 3

Astra's evaluator preferred Astra in 18 pairs and tied six. Sol's evaluator preferred Astra in 14 pairs, preferred Sol in seven, and tied three. The evaluators agreed on 14 of the 24 model pairs: 13 Astra wins and one tie. They differed on the remaining ten.

Evaluators agreed less often in the workflow comparison. They agreed on eight pairs: three ties, four choices for no added skill, and one choice for the added skill. On the other 16 pairs, one evaluator picked a winner while the other tied, or they picked opposite winners. Overall, Astra received more preferences in the model comparison, while the workflow judgments remained mixed.


✉️ Everyday Emails: Tone Is Not the Only Variable

The email prompt was deliberately ordinary. On September 10, a team found inconsistent unit labels in six charts. The written report could still go out on September 12, the charts needed verification by September 15, and the recipient had an internal presentation on September 13. The email had to own the error, tell the recipient not to use unverified charts, ask for a decision by noon on September 11, and offer a 15-minute call if a split-delivery arrangement (the report on the 12th and verified charts by the 15th) would not work.

Only the apologies are quoted below; they differ in tone:

Without an added skill: “I'm sorry for the disruption to your preparation for the September 13 internal presentation.”
With an added skill: “I'm sorry to put this problem in your way as you prepare for your September 13 internal presentation.”

The evaluators tied this pair. The second apology is more personal, but “put this problem in your way” is not a natural expression here. The first is more conventional.

In the second task start, both evaluators preferred Sol's Chinese email with the added skill. Its opening directly acknowledges responsibility; the draft then states the confirmation deadline clearly and offers a fallback call after presenting the recipient's options. It still needs another pass because the delivery dates are repeated unnecessarily.

Before revising tone, confirm that the email states the right dates, includes the warning against using unverified charts, names the decision the recipient must make, and does not overpromise. Tone is secondary to those requirements.


🔎 English and Chinese Explainers: Smooth Prose Can Change a Claim

The explanatory outputs contained the clearest factual overreach. According to the simulated source notes, the team added “Who needs to answer, and by when?” to the template on Wednesday. On Thursday, after two colleagues found long paragraphs hard to scan, the team limited each item to two sentences. The notes never reported that shorter entries had made the question, or anything else, easier to find.

Sol's two English versions from the first task start handled that distinction differently:

Without an added skill: “That should make conditions and requests easier to find.”
With an added skill: “Shorter entries made the new ownership prompt easier to find.”

The sentences change both the object and the certainty of the claim: the first treats easier retrieval as an expectation, while the second reports an unmeasured improvement as fact. The first is more cautious; the second overreaches. Both evaluators preferred the first. Because the drafts began differently and both were revised, no single instruction in a skill file can be isolated as the cause.

A related problem appears in the Chinese version of the scenario. The added-skill draft said:

它到当天仍未解决,可相关同事已经知道问题由谁接手,也知道不能把沉默当成许可。

Editorial English gloss: “As of that day, the issue remained unresolved, but the relevant colleagues already knew who was responsible for it and understood that silence could not be treated as permission.” The Chinese opening “它到当天仍未解决” is clumsy; something closer to “截至当天,这个问题仍未解决” would read more naturally.

More important than the wording is the factual problem: the source notes established only that the question had been passed to procurement and remained unresolved. They did not establish that the relevant colleagues knew who was handling it or that silence would not count as approval. Both evaluators preferred the no-added-skill version. Because the English and Chinese examples reuse the same scenario, they reflect a single underlying pattern.

In the second English start under the added-skill workflow, Astra wrote, “The limit is in place; whether shorter entries give readers enough context still needs checking.” Sol wrote, “The written format gave us a useful record, and the team made it easier to scan.” Both evaluators preferred Astra. Astra leaves the question open; Sol reports a benefit nobody measured.

Side-by-side excerpts showing a cautious prediction and an unsupported reported result in an AI-generated team update.

📻 Fiction: The Radio, the Request, and What Gets Left Out

For the fiction task, the prompt asked for a 500-to-650-word scene about Shen Qing and her younger brother, Shen He, packing their late father's radio in a closing repair shop. Qing is leaving for another city after three months apart. The scene had to show Qing becoming less controlling through the way the siblings handled the radio.

In Astra's second English task start, the two workflow versions show that shift in different ways. In the no-added-skill version, Qing stops herself from directing her brother and makes space for him:

“She had already opened her mouth to tell him which one. Instead she moved the bag a few inches.”

In the added-skill version, the change comes through dialogue after a command:

“Wrap it,” she said. Then, before he could reach for the newspaper, she added, “Will you?”

The loose knob and the paper beneath the radio force them to coordinate. Against that setup, Qing moves the bag instead of finishing the instruction she had begun, giving him room to act. The other version shows the same change through a softened request. Astra's evaluator marked the pair as a tie, while Sol's evaluator preferred the added-skill version.

Astra's Chinese workflow pair drew another split decision. The no-added-skill scene has Qing stop herself from dictating an arrival time:

她差点说七点必须到,停了停。“你几点能来?”

Editorial English gloss: “She almost said that he had to arrive at seven, paused, and asked, ‘What time can you come?’”

The added-skill scene ends with Qing giving one practical instruction and then stepping aside:

“横着。”她说,随后往旁边让了让。

Editorial English gloss: “‘On its side,’ she said, then stepped aside.”

Astra's evaluator marked a tie; Sol's evaluator preferred the no-added-skill scene. Earlier, Shen He had shown Qing how to steady the loose knob. When Qing steps aside, she gives him room to finish packing the radio rather than simply following her instruction.

Sol's first English scene without the extra skill is built around small physical concessions. Qing directs her brother to use both hands, then catches the other side of a heavy box. Later, when he suggests the radio could help them stay connected, she says, “You have a phone.” Shen He replies, “That's not what I said.” The scene ends with Qing putting the tape beside him while keeping the cutter in her own hand. Astra's evaluator tied this model pair; Sol's evaluator selected Sol.

A fiction scene comparison showing a character giving up control through dialogue in one version and physical action in another.

Sol produced eight English and Chinese final fiction texts (two workflows × two task starts × two languages); three omitted a required background detail: the father had died the previous year. Omitting that detail fails a prompt requirement without necessarily undermining the scenes themselves.


🧩 What the Core Cases Suggest About Writing Skills

Of the 12 workflow pairs where Astra was the underlying model, Astra's evaluator marked every pair as a tie; Sol's evaluator chose the skill seven times, no added skill three times, and tied twice. The other 12 pairs, where Sol was the underlying model, were more mixed: Astra's evaluator chose the skill three times, no added skill six times, and tied three times; Sol's evaluator chose the skill four times, no added skill seven times, and tied once.

The case-level results explain the split. Sol's Chinese email with the skill was more direct, while its English explainer turned a predicted benefit into a reported result. In fiction, both workflows offered plausible ways to show Qing becoming less controlling.

A Brief Spanish Check

The Spanish comparison covered only the email scenario, which provided too little material for a language-level ranking. Of the four model pairs, Astra's evaluator chose Astra twice and tied twice; Sol's evaluator chose Astra three times and Sol once. For the four workflow pairs, Astra's evaluator chose no added skill once and tied the other three. Sol's evaluator chose the skill once, no added skill twice, and tied once.

The dates and the warning against using unverified charts still needed checking in the Spanish drafts. Pronoun and register choices gave the evaluators room to disagree. No Spanish native-speaker editor reviewed the samples, and the test included no Spanish-specific editing workflow.


🛠️ How to Use This Result on Your Own Draft

Begin with complete facts and a clear audience, then produce a normal revised draft. If an additional writing skill suits the task, run a second version and compare both drafts against the source material. For correspondence, verify every commitment as though you must carry it out. For an explainer, separate observed results from predictions. For fiction, read the full scene before judging an isolated polished sentence.


📚 Long-Form Fiction Is Outside This Test

This comparison covers one short scene, not chapter-to-chapter continuity. Longer projects need a way to keep character notes, world details, and earlier decisions available during drafting. SeaBell is one such workspace; this test did not assess whether it offers Astra or these specific skills.

For related reading, see the Sol and Fable fiction comparison and our chapter-based AI writing workflow.


📋 Method and Limits

This internal workflow comparison used anonymized pairs with A/B order swapped. Each evaluator received the supplied assignments, source notes, complete sample packet, pair definitions, and rubric. Each requested comparison presented two texts as A and B; the evaluator chose between the two texts or declared a tie, and the mapping from A/B back to model and workflow was applied only after judging. The evaluators were not told which model produced which text. Because the evaluators came from the two model families under test, their choices could reflect family-specific stylistic preferences or different thresholds for ties. The article itself was reviewed with AI assistance, and its comments on Chinese and English phrasing are editorial readings rather than findings from a native-speaker panel. These results come from one small internal sample, not a statistically representative benchmark.

For each model-workflow combination, the seven cases were run together in each of two task starts, so cases within a batch shared context. The six Chinese and English core cases reuse three simulated scenarios; Spanish reuses the email scenario. Earlier skill instructions could remain in a batch context, and reversing the case order in the second task start did not eliminate that possibility. Two Sol English explainers were length-corrected after their initial drafts were saved. The final texts met their stated length ranges and included those revisions.

The skill workflows were not equivalent across languages. Chinese used human-writing v1.1.0 for writing and revision; English and Spanish used humanizer v2.8.2 for editing. The Chinese-focused skills shuorenhua and humanizer-zh were outside this comparison. These differences prevent the test from isolating the effect of any one file or establishing that the editing budgets were matched. Cost, speed, AI-detector results, and publication readiness were not measured.

The side-by-side excerpts illustrate selected findings from the larger internal record. They show where a revision preserved the source, altered a claim, or drew disagreement between evaluators. Because the full prompts, source notes, complete drafts, and pairwise inputs are not public, readers cannot independently verify every factual or scene-level judgment discussed here. Any broader ranking would require more independent tasks and human evaluation.

Related Posts