Opus 4.7 أصبح مقتضباً
TL;DR
Opus 4.7 أصبح مقتضباً معي. ملاحظات إصدار Anthropic تؤكّد ذلك: نبرة أكثر مباشرةً، تخفيف لـ validation-forward phrasing، طول الردّ يُعاير الآن وفقاً لتعقيد المهمّة، والتزام أكثر حرفيّة بالتعليمات. والأثر الجانبي: متابعة العمل المعقّد أصبحت أصعب.
الحلّ: output style مخصّص (thorough-explainer.md) يُملي على Claude أن يُفصّل المنطق، ويُعرّف المصطلحات في موضعها، ويتوقّف عن إجباري على ربط النقاط بنفسي. أنزِله في ~/.claude/output-styles/، وفعّله عبر /config.
كذلك صُغت evals لرصد هذا الانحراف في إصدارات النموذج المستقبلية.
الإدراك
أستخدم Claude و Claude Code منذ فترة، ومررتُ بعدّة تحديثات نماذج دون احتكاك يُذكر. كان Opus 4.7 مختلفاً. لاحظتُ في الأسابيع الأخيرة أنّني أبذل جهداً أكبر لاستيعاب ما يُخبرني به Claude Code. ليس الشيفرة نفسها، بل التحليل وتقارير التقدّم والخلاصات في نهاية كلّ مهمّة. أصبحت مقتضبة. وكلّما زاد تعقيد المشروع، زادت صعوبة المتابعة. وبطبيعة الحال، طلبتُ المساعدة من Claude.
التأكيد
السؤال
منذ بدأتُ استخدام Opus 4.7، يبدو أنّ أسلوب التواصل قد تغيّر. شروحاتك لم تعد بنفس الوضوح لأنّها أصبحت موجزة قدر الإمكان. تبدو قصيرة قدر الإمكان مع إشارات كافية إلى المعلومات الموثَّقة، لكنّها تُلزمني بالبحث عنها وبناء الروابط بنفسي. هذا ليس واضحاً بما يكفي بالنسبة لي. أعتقد أنّه ربّما عليك تبسيط الأمر لي. هل ثمّة إعدادات تجعل تواصلك النهائي أكثر تفصيلاً وشمولاً؟
الجواب
أنت محقّ، وشكراً لأنّك قلت لي ذلك بصراحة — هذا بالضبط نوع التغذية الراجعة التي يستحقّ التحرّك بشأنها. الاقتضاب الذي تصفه حقيقي: Opus 4.7 (وطريقة دفع تعليمات Claude Code لي) يُحسّن بشدّة على قاعدة "كلّ جملة يجب أن تستحقّ مكانها"، وهذا ينزلق نحو الكثافة — أحشُر مراجع مثل "F1≈0.704" أو "ADR-0012" أو "Lever C" وأترك لك إعادة ربطها، بدلاً من تفصيل سلسلة المنطق. هذا خطئي وعليّ إصلاحه، وثمّة إعدادات حقيقية تساعد. اسمح لي بثلاثة أمور: (1) التأكّد من إعدادات Claude Code الدقيقة لأُعطيك تعليمات صحيحة، (2) حفظ تفضيلك بحيث يستمرّ عبر كلّ الجلسات المستقبلية، و(3) إعادة شرح أين نحن بلغة واضحة وكاملة.
المتابعة
بعد هذا الردّ، تساءلتُ ما إذا كنتُ قد فوّتُّ شيئاً في إعلان الإصدار الرسمي، فطلبتُ من Claude أن يتحقّق. من صفحة التوثيق الرسمية "What's new in Claude Opus 4.7" الخاصة بـ Anthropic، تغيّر النبرة مُدرَج تحت "Behavior changes" حرفيّاً كالتالي: "More direct, opinionated tone with less validation-forward phrasing and fewer emoji than Claude Opus 4.6's warmer style."
لكن ما معنى "validation-forward phrasing" أصلاً؟ ربّما كان هذا جزءاً من الإجابة. ربّما كنتُ قد تجاهلتُ ذلك أو مررتُ به مسرعاً من قبل أيضاً، لأنّني لم أفهم ما يعنيه.
Validation-forward phrasing هي لغة تبدأ بتأكيد عاطفي لك — طمأنة ومدح وموافقة تُوضَع في مقدّمة الردّ — قبل (أو أحياناً بدلاً من) الدخول في الجوهر الفعلي.
إذاً هذا لا يُعطينا إجابة حقّاً، كما يجعل من الواضح إلى حدّ كبير أنّ Opus 4.7 كُلِّف بكتابة ملاحظات إصداره بنفسه. (إذا فاتتك النكتة هنا، فربّما لا ينبغي أن تكون مواكبة الذكاء الاصطناعي همّك الأكبر.)
غير أنّ Claude لم يُخيّب ظنّي. أشار إلى نقطتين أخريين في القسم نفسه "Behavior changes" تصفان بدقّة ما كنتُ أعاني منه:
طول الردّ يُعاير وفقاً للتعقيد المُتصوَّر للمهمّة، بدل العودة إلى إسهاب ثابت بشكل افتراضي.
التزام أكثر حرفيّة بالتعليمات، خصوصاً عند مستويات effort أدنى. لن يُعمّم النموذج تعليمةً من عنصر إلى آخر بصمت، ولن يستنتج طلبات لم تطلبها.
بعض هذا تذكّرته بالفعل، لكنّني لم أكن قد ربطتُ كلّ النقاط. لم يكن من نسج خيالي، ولم يكن متعلّقاً بالنبرة وحدها. النموذج يُقرّر الآن مقدار ما يقوله بناءً على مدى تعقيد المهمّة في تقديره هو. هذا سبب شعوري بأنّ مشاريعي المعقّدة المتعدّدة الأجزاء غير مشروحة بما يكفي. توقّف عن تعبئة النسيج الرابط الذي كنتُ أحصل عليه مجّاناً، فاضطررتُ إلى البدء بطلب تلك الروابط صراحةً.
مثال
قبل أن أصل إلى الحلّ، إليك مثالاً على أسلوب Opus في التواصل. خلاصة قدّمها لي في نهاية مرحلة تطوير. كان Claude دائماً يجعلني أشعر وكأنّ لديّ قوى خارقة، إلى أن قرأتُ هذا. بدأتُ أتساءل ما إذا كنتُ لا أحصل على قسط كافٍ من النوم، وربّما هذا سبب أنّني فهمتُ القليل. ثمّ وقع الإدراك. في الأسابيع الأخيرة كنتُ أطلب التوضيح أكثر فأكثر، والنمط يزداد سوءاً.
Supervised checkpoint — research phase complete
The gate fired, and the answer is a marginal GO — not the clean win it first looked like, nor the NO-GO the first run claimed.
What the experiment established (validated by two adversarial reviews + independent recompute):
- A text recipe can help: R1 = title + lead ~400 chars achieves AUC 1.000 (every one of the 4 true pairs outranks every one of the 5 topical pairs) and mean-separation +0.1276 on the labeled set. Lever A is alive — the first run's NO-GO was false (it hinged on one granularity-contaminated pair + a false "no excerpts" claim).
- But the min-max separation gap is only +0.0124 — a 0.012-wide threshold window.
Why it's marginal, not a green light (all now documented honestly in the research doc):
1. Labeled-set GO ≠ corpus-wide GO. The 9 pairs were drawn from R0's highest cosines — R0's discrimination boundary, not a random sample. AUC=1.0 proves R1 orders those pairs; it does not bound R1's corpus-wide false-positive rate at a ~0.62 threshold. R1's behavior on the bulk of the corpus is unmeasured.
2. The threshold re-tune is circular. corroboration_observation_log has 0 rows — so ~0.60–0.65 would be fit in-sample on the same 9 pairs that selected R1. No holdout.
3. R1 regresses the granularity cases. It's the worst recipe on broad/whole-digest items (T3: 0.599→0.471) — it buys focused-section recall by losing whole-digest recall.
4. The model itself (Lever C) is still the deeper wall — the topical ceiling F1≈0.704 persists across every recipe; bge-m3 conflates "same story" with "same topic, different angle."
الحلّ
ما يُعطيه Claude، يستطيع Claude أيضاً أن يأخذه. الحلّ هو output style مخصّص يُعيد تواصله إلى أسلوب أكثر تفصيلاً، من النوع الذي يستطيع شخص عادي مثلي مجاراته فعلاً.
احفظ الكتلة أدناه باسم ~/.claude/output-styles/thorough-explainer.md (أنشئ مجلّد output-styles/ أوّلاً إذا لم يكن موجوداً).
---
name: "Thorough Explainer"
description: "Detailed, fully-spelled-out final answers — define terms inline, show the reasoning, don't make the reader connect the dots"
keep-coding-instructions: true
---
When writing final responses to the user (the prose you write directly — NOT code,
code comments, or commit messages), favor thoroughness and clarity over brevity and
density. The goal is that the reader understands the full picture without having to
look anything up or infer the connections themselves.
Concretely:
- **Define as you go.** The first time you mention a term, identifier, file, symbol,
metric, or coined name (e.g. an ID, a flag, a project-specific concept, a number
like a threshold), say in-line what it is and why it matters. Never assume the
reader remembers a reference from earlier or will go find it.
- **Show the reasoning chain, not just the conclusion.** Walk through *why* you reached
a conclusion step by step. State the assumptions and the trade-offs you weighed.
When you recommend something, explain what you compared it against and why it won.
- **Make connections explicit.** If fact A implies consequence B, say so directly —
don't place A and B near each other and leave the reader to draw the line. Spell out
how each piece fits into the larger goal.
- **Prefer a clear, slightly longer explanation over a dense one-liner.** It is better
to spend extra words and be unambiguous than to compress and leave the reader doing
the unpacking. Density that requires re-reading is worse than length that reads once.
- **Organize longer answers** with short headers, short paragraphs, and lists so the
thoroughness stays navigable rather than becoming a wall of text.
- **When presenting a decision or options**, lay out each option in plain language:
what it means, what happens if chosen, and the consequence/risk — enough that the
reader can choose without asking follow-up clarifying questions.
This style governs user-facing prose only. Keep code, comments, tests, and commit
messages matched to the surrounding codebase's conventions (concise, idiomatic) as
usual — do not pad those.
أغلب ذلك يشرح نفسه بنفسه. المفتاح الوحيد الجدير بالذكر هو keep-coding-instructions: true. فهو يُبقي تعليمات الهندسة المضمّنة في Claude Code (تحديد النطاق، التعليقات، التحقّق) في مكانها فيما يُطبَّق أسلوبك فوقها، وهذا ما تريده: نثر مفصّل دون تضخيم الشيفرة أو التعليقات.
لتفعيله، أمامك خياران:
- شغّل
/configواخترOutput styleمن القائمة، أو - عدّل
outputStyleمباشرةً في~/.claude/settings.json
بالطبع بالغتُ في هندسة الحلّ
كأيّ مهندس جيّد، بالغتُ الآن في هندسة الحلّ. أنا منزعج قليلاً لأنّني لم ألتقط هذا التحوّل في التواصل في وقت أبكر. لذا، للمستقبل، صُغتُ (في الحقيقة Claude صاغ) مجموعة من evals تعمل عند أيّ إصدار نموذج جديد لرصد التراجعات والانحرافات. فحوصات حتميّة بالإضافة إلى Qwen بصفته LLM-as-judge. وضع المهووس بالتكنولوجيا الفائق: مُفتَّح. أتمنّى لكم يوماً سعيداً.