Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Harness tối thiểu: một task, một phép kiểm, một đường tiếp tục

Câu hỏi bài này trả lời: cần những gì quanh AI coding agent để phiên bắt đầu biết scope, phiên kết thúc có evidence và lỗi không bị biến thành “xong”?

Cần biết trước: tiêu chí hoàn tất và checkpoint. Lab dùng Python 3 stdlib (đã kiểm 3.14/macOS arm64), không cần framework hay tài khoản LLM.

Harness ở đây là môi trường quanh việc làm: instructions cho biết cách làm việc; scope/criteria định nghĩa kết quả; verifier đọc hành vi; checkpoint giữ bằng chứng và bước tiếp; lifecycle cho biết khi nào kiểm. Building effective agents khuyến nghị bắt đầu đơn giản, thêm độ phức tạp khi có ích quan sát được. Bài áp dụng điều đó bằng một runner cục bộ, không xây một platform agent.

Kernel nhỏ và ranh giới quyền

PhầnFile trong labVai trò
InstructionsINSTRUCTIONS.mdNgười/agent đọc mục tiêu và ranh giới
Scope/criteriatask.json và fixture S25Chỉ sửa exporter; expected có trước code
Verificationverify_behavior.pySo CSV đọc lại với giá trị độc lập
Statecheckpoint.json, verify.logTrạng thái hiện tại, snapshot, next step
Lifecyclerun.pySTART kiểm đầu vào → verify → ghi checkpoint

Agent hoặc người sửa code giữa hai lần chạy runner. Runner không tự repair, không lặp vô hạn đến xanh, không gọi submit/deploy. Instructions không cấp quyền OS: thật sự giới hạn file, command, network và credential phải do sandbox/tool/runtime kiểm soát. Runner lab chạy code Python đã tin cậy; nó không phải sandbox chống verifier ác ý.

Tạo thư mục trống, chép export_csv.py, fixtures.py, verify_behavior.py từ lab S25. Giữ bản no_header_when_empty đang lỗi R5. Checkpoint bài trước không cần chép: runner này tạo evidence hiện tại từ đầu.

# Cách làm task CSV

Đọc task.json và các file hiện tại trước khi sửa.
Chỉ sửa export_csv.py; giữ fixture, verifier và R1–R6.
Chạy python3 run.py sau mỗi thay đổi nhỏ.
Blocked: xử lý điều thiếu hoặc hỏi người giao việc, không tự invent expected.
Failed: sửa hành vi theo finding, không nới tiêu chí.
Error: sửa môi trường/verifier theo lỗi thật trước khi dùng kết quả.
Verified local: review diff và phần ngoài phạm vi trước khi đóng task.
Không submit/deploy bằng runner; quyền thực thi theo môi trường của nhóm.
{
  "id": "csv-header",
  "goal": "CSV đúng R1–R6; danh sách rỗng vẫn có header",
  "edit": ["export_csv.py"],
  "criteria": ["R1", "R2", "R3", "R4", "R5", "R6"],
  "expected": "Đọc lại CSV bằng parser: header và từng dòng bằng CASES/HEADER trong fixtures.py",
  "verify": ["verify_behavior.py", "no_header_when_empty"]
}

Kiểm runner trước khi tin runner

Viết probe trước run.py. Nó sửa file riêng trong thư mục lab để đưa lỗi mẫu vào, rồi khôi phục; không dùng checkout hoặc fixture thật của ứng dụng. Verifier rỗng exit 0 cũng phải bị từ chối: thiếu test không được biến thành “xanh”.

import json
import subprocess
import sys
from pathlib import Path

exporter = Path("export_csv.py")
test = Path("verify_behavior.py")
task_file = Path("task.json")
original = exporter.read_text(encoding="utf-8")
test_original = test.read_text(encoding="utf-8")
task_original = task_file.read_text(encoding="utf-8")
old = "        if users:\n            writer.writerow(HEADER)"
assert original.count(old) == 1
fixed = original.replace(old, "        writer.writerow(HEADER)")


def expect(label, status, code):
    names = ("export_csv.py", "fixtures.py", "verify_behavior.py", "task.json", "INSTRUCTIONS.md")
    before = {name: Path(name).read_bytes() if Path(name).is_file() else None for name in names}
    result = subprocess.run([sys.executable, "run.py"], capture_output=True,
                            text=True, timeout=25, check=False)
    after = {name: Path(name).read_bytes() if Path(name).is_file() else None for name in names}
    assert after == before, (label, "runner changed inputs")
    assert result.returncode == code, (label, result.returncode, result.stderr)
    state = json.loads(Path("checkpoint.json").read_text(encoding="utf-8"))
    assert state["status"] == status, (label, state["status"])
    assert state["passes"] == (status == "verified_local")
    assert state["done"] is False and state["next_step"]
    assert exporter.read_text(encoding="utf-8") == (original if label == "behavior" else fixed)
    print(f"{label}: {status} exit={code} done=false")


try:
    expect("behavior", "failed", 1)
    exporter.write_text(fixed, encoding="utf-8")
    expect("pass", "verified_local", 0)
    test.unlink()
    expect("missing-test", "blocked", 2)
    test.write_text("raise RuntimeError('fake verifier fault')\n", encoding="utf-8")
    expect("verifier-error", "error", 3)
    test.write_text("print('no checks')\n", encoding="utf-8")
    expect("empty-test", "error", 3)
    test.write_text(test_original, encoding="utf-8")
    task = json.loads(task_original)
    task["expected"] = ""
    task_file.write_text(json.dumps(task), encoding="utf-8")
    expect("missing-expected", "blocked", 2)
    task_file.write_text(task_original, encoding="utf-8")
    expect("resume", "verified_local", 0)
finally:
    exporter.write_text(original, encoding="utf-8")
    test.write_text(test_original, encoding="utf-8")
    task_file.write_text(task_original, encoding="utf-8")
print("restored inputs; checkpoint requires revalidate")

Chạy khi runner chưa tồn tại: probe phải đỏ. Đây là kiểm trước implementation, không phải lỗi cần bỏ qua trên ứng dụng thật.

python3 probe-harness.py

Runner: phân biệt chưa đủ dữ kiện, test đỏ và test hỏng

import hashlib
import json
import subprocess
import sys
from pathlib import Path

FILES = ("INSTRUCTIONS.md", "task.json", "export_csv.py", "fixtures.py", "verify_behavior.py")
OK_LINES = {"ok   normal", "ok   special_chars", "ok   unicode", "ok   empty"}
NEXT = {"blocked": "Bổ sung đầu vào/expected từ người giao việc rồi chạy lại",
        "error": "Điều tra verifier/môi trường; chưa dùng kết quả để đóng task",
        "failed": "Sửa exporter theo finding, giữ tiêu chí và chạy lại",
        "verified_local": "Review diff và scope; action submit/deploy theo quyền riêng"}


def finish(status, reason, output="", exit_code=None):
    Path("verify.log").write_text(output, encoding="utf-8")
    state = {"schema": 1, "status": status, "reason": reason,
             "passes": status == "verified_local", "done": False,
             "next_step": NEXT[status], "python": sys.version.split()[0],
             "verification": {"exit": exit_code, "log": "verify.log"},
             "snapshot": {name: hashlib.sha256(Path(name).read_bytes()).hexdigest()
                          for name in FILES if Path(name).is_file()}}
    Path("checkpoint.json").write_text(json.dumps(state, ensure_ascii=False, indent=2),
                                      encoding="utf-8")
    print(f"{status}: {reason}; done=false")
    sys.exit({"verified_local": 0, "failed": 1, "blocked": 2, "error": 3}[status])


print("START: inspect inputs")
if any(not Path(name).is_file() or Path(name).is_symlink() for name in FILES):
    finish("blocked", "missing or unsafe input")
try:
    task = json.loads(Path("task.json").read_text(encoding="utf-8"))
except (ValueError, OSError):
    finish("blocked", "invalid task input")
if (not task.get("goal") or not task.get("expected")
        or task.get("criteria") != ["R1", "R2", "R3", "R4", "R5", "R6"]
        or task.get("edit") != ["export_csv.py"]
        or task.get("verify") != ["verify_behavior.py", "no_header_when_empty"]):
    finish("blocked", "missing or unsupported task contract")
print("VERIFY: run fixed command")
try:
    result = subprocess.run([sys.executable, *task["verify"]], capture_output=True,
                            text=True, timeout=20, check=False)
except (OSError, subprocess.TimeoutExpired):
    finish("error", "verifier could not complete")
output = result.stdout + result.stderr
lines = set(result.stdout.splitlines())
if result.returncode == 0 and OK_LINES <= lines:
    finish("verified_local", "four behavior cases passed", output, result.returncode)
if result.returncode == 1 and any(line.startswith("FAIL ") for line in lines):
    finish("failed", "behavior differs from expected", output, result.returncode)
finish("error", "verifier crashed or missing result protocol", output, result.returncode)

Quy ước output bốn ca là contract nhỏ của lab, không dùng để suy mọi test suite đáng tin. Một verifier cố ý in bốn dòng mà không kiểm gì vẫn lách được; review/mutation test và bảo vệ verifier thuộc môi trường tin cậy là phần còn lại. expected không rỗng cũng chưa chứng minh đặc tả đúng ý người dùng.

python3 probe-harness.py
behavior: failed exit=1 done=false
pass: verified_local exit=0 done=false
missing-test: blocked exit=2 done=false
verifier-error: error exit=3 done=false
empty-test: error exit=3 done=false
missing-expected: blocked exit=2 done=false
resume: verified_local exit=0 done=false
restored inputs; checkpoint requires revalidate

Probe đã khôi phục source lỗi để lần chạy khác lặp được. Checkpoint xanh cuối probe vì thế không còn mô tả source hiện tại. Chạy lại runner phải đỏ:

python3 run.py

Sửa đúng nhánh header rồi tiếp tục. Runner chỉ kiểm/ghi artifact; thao tác sửa là bước work riêng.

python3 - <<'PY'
from pathlib import Path
path = Path("export_csv.py")
source = path.read_text(encoding="utf-8")
old = "        if users:\n            writer.writerow(HEADER)"
assert source.count(old) == 1
path.write_text(source.replace(old, "        writer.writerow(HEADER)"), encoding="utf-8")
PY
python3 run.py

Khi nào thêm, khi nào bỏ thành phần?

Một task nhỏ có thể chỉ cần instructions ngắn, verifier có ý nghĩa và một checkpoint. Chỉ thêm wrapper khi có lỗi lặp lại mà wrapper bắt được. Harness cho agent chạy dài nhấn mạnh artifact bàn giao và kiểm đầu phiên; case này thử được đường lỗi ở runner, chưa đánh giá độ tin cậy của model.

Muốn thử bỏ một component, dùng tập task/ca lỗi đã chốt, cùng model/tool/quyền và điều kiện; so bản có/không component trên nhiều lượt. Đo completion đúng theo criteria, resume đúng, false block, thời gian và chi phí. Giữ critical verifier/quyền thật trong môi trường cô lập khi thử; không tắt bảo vệ production để benchmark. Nếu không có cải thiện quan sát được, đơn giản hóa rồi kiểm lại các ca lỗi. Đây là kế hoạch đo chưa thực hiện ở lab, không có tỷ lệ cải thiện để công bố.

Giới hạn: runner không khóa file, không bảo vệ criteria khỏi agent sửa, không sandbox import Python hay giữ credential; hash không chữ ký. Thư mục lab chỉ có dữ liệu giả. Với hệ thống thật, evidence chứa thông tin riêng phải được lưu và cấp quyền đúng. Không biến hook giả hoặc câu chữ “cấm” thành bằng chứng rằng hành động đã bị chặn ở runtime.

Học tiếp checkpoint và stale evidence, review code AI. Nguồn Anthropic đọc ngày 2026-10-03; ví dụ runner là thiết kế riêng.