この資料は英語です
リファレンス(文法・形式・生成物・エージェント向けの手順)は、第一の読み手がエージェントなので英語で書いています。日本語で読めるのは、ホーム・インストール・表を書く・何を証明するか・生成して呼ぶ・突き合わせと再生・例で見る・自分のルールが入るか、そして診断コードの台帳と互換性です。
Targeting a language rulec does not generate
rulec gen writes Python, TypeScript, JavaScript, Rust, Ruby, PHP, Go, Swift, Java, SQL, Wasm and NumPy. This page is about the
thirteenth target — a language nobody planned for, a workflow engine's expression language, a
spreadsheet formula, a database.
The short answer: you do not have to modify rulec, and you do not have to give up the
evidence. A rule already publishes everything a generator needs, and rulec verify will
stand any process up and hold it to the rule over every generated case. The rest of this page
is that loop, run end to end against SQL by hand — as it was done before SQL had a backend of
its own (rulec gen writes one now, in a different shape); the loop is the same for any
target.
The two paths
| Generating from outside | A built-in backend | |
|---|---|---|
| Where the code lives | your own generator, any language | src/codegen.rs |
| Changes to rulec | none | a Lang variant and a file of its own, about 1,100 lines |
| Reaches | anything you can run | anything with a runner |
Gets rulec gen, rulec test, rulec api |
no | yes |
| Evidence | rulec verify over the generated cases |
the agreement test in the suite |
Generating from outside is the general answer and the one to reach for first. Building a backend in is what a target earns after it has been used enough to be worth carrying; the last section says what that costs.
What makes a target trustworthy
The emitter is the easy half. Whatever the target, a rule is a handful of conditions and a
handful of amounts, and turning those into CASE or if or a formula is mechanical.
The half that matters is the evidence, because generated code outside the comparison is
outside the claim this tool makes. A row transcribed with < where the rule says <=, a
rounding applied in the wrong place, a rate read as a whole number — none of those look wrong
on the page. They look wrong only when something runs both implementations on the same inputs
and compares the answers.
So the order below is: take the inventory, emit, and then make the target answer questions. The last step is not optional decoration; it is the reason the first two are worth doing.
Generating from outside, step by step
The worked example is tests/corpus/ec261.rule — Article 7 of
Regulation (EC) No 261/2004, three stacked tables and a rate — targeting SQLite.
1. Take the inventory
One JSON object per language with the module and function names, the parameters in order with
their brands, units and ranges, the outputs with their rounding, the enum members, and the
errors the code can raise. Read the names from here rather than from the .rule source: the
inventory already applies the ASCII aliases and the per-language spelling rules, so a name you
take from it is the name the other implementations use.
For the table logic itself, read the rule's own structure with rulec check --format json or
parse the .rule; it is a fixed grammar, described in reference.md.
2. Take the wire
The JSON Schema of the in and out objects. Everything on the wire is an integer in the
canonical unit — no decimals anywhere. Two consequences bite generators:
- A rate travels as a count of steps. A
rate[step 50%]column carries1for 50% and2for 100%. Multiply by it and the product is held in halves; divide that out at the end. - Rounding is the last thing that happens, on the value in its own scale, and the
direction is declared in the rule.
round downis toward zero, not toward negative infinity:-4.8 EURrounds down to-4 EUR. generated-code.md has the five modes.
There is a second spelling of the same wire. rulec gen also writes the rule's contract
as a .proto — the request, the answer with the rows that decided it, and one method
(generated-code.md, the rule as a Connect service).
A target whose language has a protobuf plugin can take that file and let the plugin write the
messages: the names, the types and the enum values are then the same on both sides by
construction, and what you write is the call into your own code. Only Python gets a
generated service, but the .proto is not Python's.
3. Emit whatever your target needs
Nothing here is rulec's business — this is your generator. For SQL the shape is a chain of CTEs, one per table, each adding its output column:
WITH input AS (
SELECT :distance AS distance, :intra_eu AS intra_eu, :delay AS delay
),
band_of AS (
SELECT *, CASE
WHEN distance <= 1500 THEN 'short' -- row 1
WHEN distance > 1500 AND intra_eu THEN 'medium' -- row 2
WHEN distance > 1500 AND distance <= 3500 AND NOT intra_eu THEN 'medium' -- row 3
WHEN distance > 3500 AND NOT intra_eu THEN 'long' -- row 4
END AS band FROM input
),
amount AS (
SELECT *, CASE band
WHEN 'short' THEN 250 -- row 1
WHEN 'medium' THEN 400 -- row 2
WHEN 'long' THEN 600 -- row 3
END AS base FROM band_of
),
reduction AS (
-- factor is a rate at step 50%, so it travels as a count of steps: 1 = 50%.
SELECT *, CASE
WHEN distance <= 1500 AND delay <= 2 THEN 1 -- row 1
WHEN distance <= 1500 AND delay > 2 THEN 2 -- row 2
WHEN distance > 1500 AND intra_eu AND delay <= 3 THEN 1 -- row 3
WHEN distance > 1500 AND intra_eu AND delay > 3 THEN 2 -- row 4
WHEN distance > 1500 AND distance <= 3500 AND NOT intra_eu AND delay <= 3 THEN 1 -- row 5
WHEN distance > 1500 AND distance <= 3500 AND NOT intra_eu AND delay > 3 THEN 2 -- row 6
WHEN distance > 3500 AND NOT intra_eu AND delay <= 4 THEN 1 -- row 7
WHEN distance > 3500 AND NOT intra_eu AND delay > 4 THEN 2 -- row 8
END AS factor FROM amount
)
-- result compensation = base × factor, held in halves of a EUR, then
-- round down(1EUR): toward zero, which is ABS/g*g carrying the sign back.
SELECT (CASE WHEN base * factor < 0 THEN -1 ELSE 1 END)
* (ABS(base * factor) / 2 * 2) / 2 AS compensation
FROM reduction;
Keep the row comments. Reading the target side by side with the rule is how a person checks it, and it is what makes a mismatch report legible when one arrives — the built-in backends do the same thing for the same reason.
4. Wrap it so it answers questions
A 20-to-30 line template that exchanges JSON Lines over stdin/stdout: a handshake, then one line in and one line out per case. formats.md has the protocol. Filled in for the SQL above, using nothing but the Python standard library:
"""Wrap the generated SQL so `rulec verify` can hold it to the rule."""
import json, sqlite3, sys, pathlib
SQL = pathlib.Path(__file__).with_name("ec261.sql").read_text()
db = sqlite3.connect(":memory:")
sys.stdin.readline() # handshake
print(json.dumps({"ok": True, "impl": "sqlite3/ec261.sql"}), flush=True)
for line in sys.stdin:
line = line.strip()
if not line:
continue
req = json.loads(line)
d = req["in"]
row = db.execute(SQL, {"distance": d["distance"],
"intra_eu": int(d["intra_eu"]),
"delay": d["delay"]}).fetchone()
print(json.dumps({"id": req["id"], "out": {"compensation": row[0]}}), flush=True)
The adapter is the only part that has to speak the protocol, and it can be written in whatever language can start your target. A target that is not a program at all — a spreadsheet, a rules engine hosted somewhere — needs an adapter that drives it; that is the whole of the work, and the section after next is about what to do when it cannot be done.
5. Hold it to the rule
$ rulec verify rules/ec261.rule --adapter python3 adapter.py
Compared 94 / matched 94 (100.000%)
Counterpart: sqlite3/ec261.sql
No mismatches.
Almost none of those 94 cases is hand-written. rulec builds them from the rule's own
boundaries — both sides of every comparison, one case per row, and the pairs where one row
shadows another — and adds the rule's own examples, which are the six a person did write.
It is the same suite rulec gen writes next to the generated code and rulec coverage audits.
You can also take it directly with rulec vectors.
Exit code 0 is agreement, 1 is mismatches, 2 is a bad argument or an adapter that would not
start. --format json gives the same result as data.
6. Check that the evidence bites
A check that has never failed has not been shown to work. Break one thing and look. Article 7(2)(c) allows four hours for the longest band; here the generator wrote three, which is the threshold from the band above and exactly the kind of slip a transcription makes:
$ rulec verify rules/ec261.rule --adapter python3 adapter.py
Compared 94 / matched 92 (97.872%)
Counterpart: sqlite3/ec261.sql
Affected 2 (2.128%) amount -600
table band_of row 4 / table amount row 3 / table reduction row 7 2 records difference -300 uniform total -600
Example: delay=4, distance=3501, intra_eu=false → rule compensation=300 / legacy compensation=600
It names the rows that fired in all three tables, how much money the difference is worth, and a concrete case that shows it. That report is the deliverable: it is what you hand to whoever signs the rule off.
What the evidence covers, and what it does not
Covers. Every case rulec builds from the rule's boundaries, compared answer for answer
against the reference evaluator. rulec coverage says what that suite reaches — every row,
both sides of every boundary, every shadow pair, every computed value at two values, every rounding
tie — and it is the same suite the built-in backends are held to.
Does not cover. It is a comparison over a finite suite, not a proof of equivalence; the static proofs (completeness, overlap, units, overflow) are about the rule, and they are done before any of this and hold whatever you generate. It says nothing about whether the table matches the world, and nothing about the parts of your target the rule does not describe: where the value comes from, what happens when the input is outside the declared range, or anything the host does around the call. The declared range in particular is an obligation your target inherits — the built-in backends emit an entry guard for it, and yours should too.
When the target cannot be run at all
Some targets cannot be executed where you are: an engine you have no licence for, a spreadsheet nobody can drive from a script, a DSL whose interpreter is somebody's service. There is then no adapter to write and no comparison to run.
Generate it anyway, and say on its face that it carries no evidence. Put a header in the output that says so, name the rule and version it came from, and keep it out of any list of verified artifacts:
-- Generated from ec261 v1 by <generator>.
-- NOT VERIFIED: this target was not compared against the rule.
-- The static proofs hold for the rule; nothing here has been checked against them.
The reasoning is that the alternative is worse. Refusing to generate does not stop anyone from
needing the rule in that system — it sends them to transcribe it by hand, which is the failure
this tool exists to prevent, and it leaves no marker at all. An unverified artifact that says
it is unverified is the honest form of that situation. It is still outside the claim: do not
describe it as proved, and do not list it beside a target that passed verify.
This is a deliberate widening of the rule in DESIGN §15.13 ("a language that cannot join the agreement check does not go in"), which governs built-in backends and still does. What you generate from outside is yours, not rulec's, and the labelling is what keeps the two apart.
A built-in backend
Worth doing when a target is used often enough that regenerating it should be one command, and
when its output should be held by the suite rather than by you. It buys rulec gen,
rulec test, an entry in rulec api, and a place in the agreement test.
What it costs, measured four times. The third target, TypeScript, took about 700 lines of emitter, and the fifth, Ruby, about 500.
The tenth took 1,098 and the eleventh 1,293, both in a file of their own under
src/codegen/ — and most of what is above the earlier figures is the doc comments that say
why each decision went the way it did. Plus one row elsewhere.
That "one row" is recent. Adding Ruby meant editing twenty-two files, because the set of
backends was written out again in seven separate lists. It now lives once, in
src/backend.rs, and gen, test, api and the test suites all walk it. What is left
outside is prose, and a test holds that to the registry too: a paragraph that names three of
the targets and not the rest fails.
The pieces, in the order they are usually written:
- A
Langvariant, and its spellings inLang::spelling()— the comment marker, how two conditions are joined, how anifopens, what closes the block. Shared code reads those rather than matching on the language. Gen::<lang>()— the module: brands or their absence, enums, the rounding helpers, the entry guards, one branch per row with the row quoted in a comment.Gen::<lang>_runner()— reads the wire JSON on stdin, writes the canonical JSON on stdout. This is what lets the target join the agreement test.round_tests_<lang>()— the five rounding modes against the reference values.- An arm in
Gen::cell(), and an entry inGen::api(). - A row in
src/backend.rs: the id, the name, the toolchain, the files it writes, how to run them, and — where the files are not named after the rule's alias, as Java's are not — what it does name them. Nothing else has to be told about the language —rulec genwrites it,rulec testruns it, the agreement test compares it, and the test that holds the documents to the registry starts requiring the prose to name it.
Whatever the language's own type system can carry, carry it — and say plainly what it cannot.
Rust, Swift and Go hold the unit in the type; TypeScript brands a bigint and tsc --strict is
run over the output; Python declares a NewType that a type checker enforces and
mypy --strict is run over the output to prove it;
Ruby, PHP, JavaScript, Java and SQL cannot hold a unit at all, so there it is documented
instead, and the .rbs that ships with the Ruby module says so too; the Wasm module is the
Rust one, so the unit rides in it and the .wit states it for the wire; the NumPy plan has no
type system to lean on at all, so the unit and the scale are fields of the plan and rulec api
prints them. Carrying something
still pays where the unit cannot ride: PHP and Java declare the kind of every parameter, so
their entry guard is the one Go, Rust and Swift emit — the range alone — while the three
languages that declare nothing have to ask at the door whether a number is an integer.
The rule that decides whether a target may be built in has not moved: it must be able to join the byte-for-byte agreement check. A generated artifact the suite cannot run is outside the claim, and inside the repository that is not allowed. Outside it, generated from outside and labelled as unverified, it is.
A plugin for the host's own system
A fair question, once there are twelve targets: why not a mode that packages the output as a plugin for each one's own ecosystem — an npm package, a gem, a composer package, a Maven artifact, a WordPress plugin, a Rails engine?
The question turns out to be two questions, and they have opposite answers.
A package manifest is configuration. The package name, the version, the licence, the
namespace, the organisation's scope: none of that is in the .rule, so the tool would have to
invent it or grow a configuration file to be told it — and "no runtime and no configuration"
is the whole promise of the output. It is also the part verify cannot hold to anything: an
agreement check compares answers, not metadata. The line that is already drawn explains it:
a manifest is written only when the file will not run without one. go.mod is there
because go run needs it; a Cargo.toml was considered and dropped for the same reason in
reverse. Neither of the two newest targets needed one — a require_once and a javac
invocation are enough. Packaging a library of rules is a real job, but it is one the
library's repository does once, not one the generator does per rule.
A host's plugin shape moves. Shopify Functions changed target, entry point and input transport inside a year (below). A page that goes stale says so to anyone who reads the date; generated scaffolding that goes stale looks authoritative. That is why the boundary belongs in your code and this page holds the recipe.
What the tool does generate is a door, and a door has to meet three conditions: it speaks
the wire the tool already speaks (JSON), it needs no SDK, and rulec test can drive it over
the vectors. The MCP server, the page for people and the canonical-ABI Wasm module are the
doors that met all three. Anything that does not is a recipe — the two below are the worked
ones — plus verify, which holds whatever you build to the rule just the same.
A Wasm host: Shopify Functions
Two shapes of Wasm come out of rulec gen, and a platform takes one or the other. The
wasm/ target is a module that exports a function — call: func(input: string) -> string
in the canonical ABI, with a .wit that makes a component of it — for a host that calls
into it: a script host, an Extism plugin, a component runtime
(generated-code.md). The other shape is a WASI command that reads
one JSON document on stdin and writes one on stdout; that is what the Rust runner is, and it
compiles for wasm32-wasip1 unchanged
(generated-code.md).
A Shopify Function is on the first side, through a door of the platform's own. It is built
for wasm32-unknown-unknown, and for every extension target it exports one function that
takes no arguments and returns nothing — cart.lines.discounts.generate.run, say. It reads
the input and writes the operations through imported host calls rather than through stdio:
the Shopify Wasm API, which the shopify_function crate wraps. The module has to stay under
256 kB, the run under 11 million instructions, and the input_query that shapes the input
under 3,000 bytes. No network anywhere. None of that is a shape gen emits, and none of it
has to be: the decision itself is small — the wasm/ module for member_shipping_fee is 44 kB, built with
rustc alone — so what the budget goes on is the boundary, not the table.
The rule stays a rule: flat inputs, one decision. What the function adds is the boundary —
the twenty lines that pull the scalars the rule wants out of the input, call the generated
function, and turn its outputs into operations. Call the generated Rust module by its typed
signature (rust/<alias>.rs, the same file the wasm/ target wraps); the JSON wire of that
target would cost a parse and a serialisation against the instruction budget and buy nothing
here. "Three or more refrigerated items" and "the dearest line" are written in the rule
(count, and elements with fold), so the boundary passes the lines through and does not
decide anything. Keep it in the function's crate beside the generated module, with ordinary
tests of its own; the generated part is the part the vectors cover.
The boundary is also where the platform's own shape lives, and that shape moves: the input
query, the operations and the API version are Shopify's, on Shopify's release cycle — the
target used to be wasm32-wasip1 and the input used to arrive on stdin. That is why the
boundary belongs in the function's crate and not in this tool. function-runner runs the
built module over one input document and reports the instructions it took, which is how the
boundary earns the same kind of evidence the rule already has.
What the platform decides is out of the rule's reach. A shipping rate is not set by a
function (the delivery functions rename, reorder and hide options; rates come from a carrier
service), so a tariff like yupack_base_fee is served from the generated code behind an HTTP
endpoint, and a discount table like coupon_discount_amount is the function itself.
A spreadsheet again: Google Apps Script
rulec import xlsx is the way in from a workbook. This is the way back, and it closes the
loop: the person who owns the policy goes on working in the sheet, but the cell now calls
something with a completeness proof and a rulec doc rendering behind it.
Apps Script runs V8, so the generated JavaScript module is the module — with two things to know.
There are no ES modules. Apps Script concatenates .gs files into one global scope, and
export is a syntax error there. Strip it as a build step, never by hand, so that
gen --check still guards the file it came from:
Every integer in the module is a bigint, because a JavaScript number is exact only to
2^53 and the overflow proof (E108) is against int64. A cell gives a number and takes one, so
the wrapper converts at both ends:
/**
* The shipping fee, in yen.
* @param {string} dest
* @param {number} weight g
* @param {number} total yen
* @param {string} member
* @return {number} yen, rounded up to 10 yen
* @customfunction
*/
function SHIPPING_FEE(dest, weight, total, member) {
return Number(member_shipping_fee(String(dest), BigInt(weight), BigInt(total), String(member)));
}
BigInt(1.5) throws, which is the right answer: the rule refuses a non-integer at the door
anyway. An error thrown inside a custom function reaches the cell as #ERROR! carrying the
message, so the entry guard ends up stating its refusal where the person who typed the value
can read it. Number() on the way out would lose precision above 2^53 — a yen amount never
reaches that, and saying so is cheaper than leaving a reader to wonder. A second custom
function around member_shipping_fee_traced puts the row that fired in the next column, which is
what a sheet wants for the same reason a log line does.
Where the evidence stops. The module is the one the JavaScript pass of rulec test
proves — the sed removes a keyword and nothing else — and the wrapper is the four lines
outside it. That is the whole boundary, and it is small enough to read. If you want it inside
as well, clasp run calls a deployed function from the command line, and that is a process
rulec verify can drive over the vectors; it needs your credentials and the network, so it
belongs on your machine rather than in CI.
None of this is a backend. A .gs that nothing on the machine can run cannot ride in
rulec test, and a target outside the comparison is outside the claim — so the recipe lives
here, and gen writes what it can already prove.
Where to read next
- reference.md — the grammar, so a generator can read a
.rule - formats.md — the wire, the vectors, and the adapter protocol in full
- generated-code.md — what the built-in backends emit and guarantee
- codes.md — every diagnostic, for a generator that reports back