{"entries":[{"file_sha256":"6410a79f49a8948c0858239ee95ae88854c74469afe91ec0e22773de50fba131","kind":"documentation_input","lines":"1-20","path":"docs/spec-compiler-part-3-mem.md","roles":["applicability","context"],"text":"# Loom Compiler Part 3 Memory Frontier Lowering\n\nThis document is the memory-order source of truth for graph-local SCF to\nDataflow lowering. The concrete owner is `loom-lower-graph-memory`; it\nnormalizes supported memory leaves and recursively lowers structured graph\nregions in one traversal.\n\nThe Dataflow operation contracts remain owned by the Dataflow specifications.\nThis document defines only the compiler analysis state and the ordinary SSA\nevent network produced from it.\n\nThe resulting canonical memory actors and their explicit `ctrl` and `done`\nnetwork are canonical software semantics. Their operation contracts are owned\nby `docs/spec-dataflow-memory-consistency.md` and\n`docs/spec-dataflow-vectorization.md`. TechMapping, SpatialMapping, and\nSystemMapping may realize that network on Fabric resources, but they must not\nreconstruct missing memory order from source order, graph text order,\ntraversal, or physical placement. The downstream realization boundary is\nspecified by `docs/spec-mapping-memory.md`.","why":"Names loom-lower-graph-memory as the concrete owner of graph-local SCF to Dataflow memory-order lowering, which is the sampled stage, and states that the produced ctrl/done network is the canonical semantics the postcondition reads."},{"file_sha256":"6410a79f49a8948c0858239ee95ae88854c74469afe91ec0e22773de50fba131","kind":"documentation_input","lines":"21-46","path":"docs/spec-compiler-part-3-mem.md","roles":["input_construction","input_well_formedness"],"text":"## 1. Scope\n\nThe lowering contract covers:\n\n* scalar and fixed-ranked vector forms of canonical `dataflow.load` and\n `dataflow.store`, including the masked contiguous and gather/scatter forms\n defined by `docs/spec-dataflow-vectorization.md`;\n* canonical atomic load/store, `dataflow.atomic_rmw`,\n `dataflow.cmpxchg`, `dataflow.fence`, and volatile access contracts defined\n by `docs/spec-dataflow-memory-consistency.md`;\n* normalized scalar `memref.load` and `memref.store` leaves over a canonical\n linear memory space;\n* sequential composition;\n* arbitrary nesting of `scf.if`, source-sequential `scf.for`, and\n `scf.while`;\n* basic graph-local alias-root partitions;\n* conservative unknown accesses;\n* value, execution, write-frontier, and read-frontier projection through the\n same structured selectors;\n* pre-mutation rejection of residual `scf.parallel` and `scf.forall` that\n reach a graph without an already materialized schedule boundary.\n\nThe lowering does not select parallel width, ownership, serialization,\nunrolling, reduction order, or any other schedule policy. Those decisions\nmust be made before graph-region lowering and normalized into supported\nstructured input.","why":"Scope of accepted input: normalized scalar memref.load/store leaves over a canonical linear memory space, sequential composition, arbitrary scf.if / source-sequential scf.for / scf.while nesting, pre-mutation rejection of residual scf.parallel and scf.forall, and the rule that no schedule policy is chosen here. This fixes exactly what the grammar may sample."},{"file_sha256":"6410a79f49a8948c0858239ee95ae88854c74469afe91ec0e22773de50fba131","kind":"documentation_input","lines":"75-161","path":"docs/spec-compiler-part-3-mem.md","roles":["input_construction","input_well_formedness"],"text":"## 3. Basic Alias Partitions\n\nPartition identity is local to one `dataflow.graph` lowering run.\n\nA canonical root is found by peeling an accepted side-effect-free memref view\nuntil reaching an explicit storage or boundary root. The finalized surface\nrecognizes:\n\n* a graph memory input, whose root identity comes from its launch binding;\n* a `dataflow.memory.service` result at that binding, which preserves the root\n of its exact pointer operand while changing only the value-plane pointer into\n a memory-plane capability;\n* a fresh `memref.alloc` result, whose root is unique for each invocation;\n* a verified side-effect-free view that preserves the source root. The initial\n accepted set contains `memref.cast`; adding another view form requires one\n matching root, region, and simulator contract before admission.\n\nWhen graph publication can trace every captured memory capability to a known\nroot, an exact service rooted at a unique thread argument mechanically inherits\nthat argument's `llvm.noalias` fact. If a root is unknown, appears through more\nthan one captured capability, or does not resolve to that argument, publication\nmust omit the fact. The service result does not independently assert aliasing,\nand graph publication does not perform another alias analysis.\n\nGraph launch memory bindings require exact memref capability types. An LLVM\npointer cannot bind a graph memref through a conversion, inferred base, or\nspecial address-space-zero rule. SCF optimization may first prove and\nmaterialize a rooted memref capability plus integer offset, or it may retain\nthe pointer as a value consumed by a `PointerAddressed` memory actor together\nwith an independently bound service capability. Neither path materializes a\ngraph-body bridge. `builtin.unrealized_conversion_cast` is never a canonical\nroot, view, actor, or boundary bridge.\n\nThe Canonical Dataflow finalizer assigns one `LogicalMemoryRootRef` to each\nstatic imported-memory formal role and canonical fresh-allocation definition.\nAn imported graph memory argument does not create a competing root: its exact\n`dataflow.graph.launch` binding resolves through root-preserving views to the\nupstream static role. A fresh allocation result is the root-defining value.\nView operations remain typed structural relations and receive no root ID of\ntheir own.\n\nPersistent consumers use the closed forms owned by\n`docs/spec-compiler-part-3-dfg.md`: `LogicalMemoryViewRef`,\n`LogicalMemoryRootOrViewRef`, and `MemoryExposureRef`. This document does not\nredeclare their wire variants.\n\nThe root-local inventory resolves every admitted static view to its unique\nroot-preserving relation. Reusing one graph under different roots creates\nseparate structural view references in those root inventories rather than a\nview entity. A memory exposure identifies one launch-contextual graph memory\nresult. It describes a provided capability boundary, not a token producer or\nan addressed memory operation.\n\nThis persistent reference identifies a static software role. Runtime object\nidentity is derived separately: an import is bound through the exact launch\nand runtime memory registry, while a fresh allocation combines its static root\nreference with the graph invocation occurrence. Two imported roles may resolve\nto one runtime object through explicit alias topology without merging their\nstatic IDs. Partition identity below remains local analysis state and is not\nthe persistent root catalog.\n\nA memory input binds an established external memref capability through an\nexact graph-launch type match. An LLVM pointer never satisfies a graph memory\nport. A first-class pointer value used by a `PointerAddressed` actor resolves\nthrough the runtime object registry to one object and byte offset independently\nof the service-capability binding.\n\nDistinct graph memory inputs are conservatively may-alias unless explicit\nno-alias evidence distinguishes them. Distinct fresh allocations are\nindependent roots. The analysis does not use address ranges, affine\ndisjointness, bank identity, physical ports, or element-type compatibility to\nsplit a root.\n\n`memref.get_global`, `memref.alloca`, globals, static pointer bases, and\nunrecognized capability producers are not canonical roots. A pre-final\nanalysis may conservatively group an unresolved access while building an event\nnetwork, but finalization rejects any such residual producer rather than\ngranting it an external-memory authority.\n\nA source-origin `llvm.alloca` accepted by the Structured\n`PromoteOrderedBufferToChannel` decision is not an exception to this rule. That\ndecision must remove the complete proved allocation closure before D0; a\nresidual allocation or pointer use remains non-canonical and is rejected.\n\nAccess-to-partition membership is kept in a transient operation map before\nSCF operands are projected. Selector demuxing must not change alias identity.\nThe map is discarded after explicit event edges are emitted.","why":"Canonical alias roots and capability binding: graph memory inputs bound by exact memref type, memref.alloc, memref.cast views; memref.get_global, memref.alloca, globals, static pointer bases and LLVM pointers are not roots; distinct graph memory inputs are conservatively may-alias. The grammar therefore binds every access to a graph memref input argument."},{"file_sha256":"6410a79f49a8948c0858239ee95ae88854c74469afe91ec0e22773de50fba131","kind":"documentation_input","lines":"294-330","path":"docs/spec-compiler-part-3-mem.md","roles":["context","applicability"],"text":"## 7. Source-Sequential `scf.for`\n\n`dataflow.stream` produces `K` valid induction values and a `T^K F` loop\nselector. Index bounds are cast to the configured integer index width before\nthe stream and the induction value is cast back for source index uses.\n\nThe loop owns independent recurrence rings for:\n\n* execution permission;\n* every source iter_arg;\n* `W[p]` for each touched partition;\n* `R[p]` for each touched partition.\n\nRequired sequenced-before tails use the same condition-driven recurrence\nmechanics when they cross an iteration. This does not serialize unrelated\nplain accesses or create a persistent loop-order object.\n\nEach ring uses `dataflow.carry` under the loop selector. A matching\n`dataflow.demux` sends true-lane values into the body and the false-lane value\nto loop exit. Captured non-memory values are replayed with\n`dataflow.invariant` and projected into body phase with `dataflow.gate`.\nMemref capabilities are not replayed.\n\nThe recursively lowered body supplies all recurrence feedback values. The\nexecution feedback is the body's structural exit; memory feedback is the\nbody's resulting frontier pair and any path-live sequenced-before tails. The\nrings are independent even when a write assigns the same `done` to both memory\ncomponents.\n\nFor zero trip count, the stream emits only `F`. No body address or access\nfires. Every carry exposes its init value on the false lane, so source values,\nexecution, `W`, `R`, and `SB` transfer through the loop unchanged.\n\nNo dependence is removed because a loop appears parallelizable. Source\niteration order remains authoritative until an earlier transformation has\nmaterialized a different schedule with provenance.","why":"Section 7 defines the source-sequential scf.for lowering the sampled obligation belongs to: independent dataflow.carry recurrence rings for execution, iter_args and W[p]/R[p], the demux lanes into the body, and the rule that the lowered body supplies all recurrence feedback. This is the terminology the postcondition uses to recognize a lowered loop and its memory-order ring, and it ends with the sampled sentence at lines 327-329."},{"file_sha256":"6410a79f49a8948c0858239ee95ae88854c74469afe91ec0e22773de50fba131","kind":"documentation_input","lines":"380-402","path":"docs/spec-compiler-part-3-mem.md","roles":["input_well_formedness"],"text":"## 10. Parallel Transfer Boundary\n\nResidual `scf.parallel` or `scf.forall` is checked across every graph before\nthe pass mutates any graph. Raw or unowned parallel input fails. A fixed finite\nparallel region is accepted only when its Structured Program Candidate owns a\ntyped, verifier-proven `P[]` schedule and the recursive transfer can derive one\ncomplete frontier relation for that exact domain. The lowering must not trust\nthe mere presence of string-named attributes as proof.\n\nUntil the typed producer and verifier establish this provenance, the boundary\nfails closed. Forged, malformed, foreign-owner, or domain-mismatched\nprovenance is invalid even when the residual SCF shape is otherwise supported.\nPart 3 consumes the selected schedule; it does not choose parallel width,\nserialization, ownership, or reduction order.\n\nThe graph-region owner does not:\n\n* infer a width or ownership domain;\n* serialize the region;\n* unroll it;\n* choose reduction order;\n* use traversal order as a hidden schedule.","why":"The parallel transfer boundary: raw or unowned scf.parallel/scf.forall fails before mutation and provenance must be typed and verifier-proven. The grammar consequently never samples residual parallel SCF, so the sampled loops are exactly the source-sequential ones the claim governs."},{"file_sha256":"6410a79f49a8948c0858239ee95ae88854c74469afe91ec0e22773de50fba131","kind":"documentation_input","lines":"468-494","path":"docs/spec-compiler-part-3-mem.md","roles":["input_well_formedness"],"text":"## 12. Supported Failure Modes\n\nThe owner rejects before mutation when:\n\n* raw or unverifiably owned parallel SCF reaches a graph;\n* an effectful or unmodeled nested operation reaches a graph;\n* a residual LLVM load, store, atomicrmw, cmpxchg, fence, memcpy, memmove, or\n memset remains after\n normalization and therefore has no explicit completion event;\n* a source memory access has not been normalized to the canonical linear\n memory-space form required by its scalar or vector Dataflow actor;\n* structured control carries a memref result or memref loop state;\n* the graph entry lacks the leading `none` execution value.\n\nLLVM memcpy, memmove, and memset intrinsics are expanded into their exact\nstructured loop semantics before ownership selection. Supported LLVM\nload/store (including volatile and atomic contracts), `atomicrmw`, `cmpxchg`,\nand `fence` forms are then normalized before recursive region lowering, after\nwhich the same frontier rules apply. LLVM target-specific sync scopes without\na compiler-target owner and atomic accesses without an explicit power-of-two\nsource alignment fail closed. Every residual raw LLVM memory operation fails\nclosed. The finalized-graph gate also rejects residual\n`memref.load`/`memref.store`, `memref.get_global`, raw pointer arithmetic,\npointer-bearing operations, `builtin.unrealized_conversion_cast`, and unknown\nmemory-capability producers. An unsupported effectful operation inside a\nstructured region must likewise fail closed instead of being hoisted.","why":"Supported failure modes: residual raw LLVM memory operations, unnormalized accesses, memref-carrying structured control, missing leading none execution value, residual memref.load/store at the finalized gate, unrealized conversion casts. Every sampled graph avoids all of these so the subject accepts the input and the claim is exercised rather than the reject path."},{"file_sha256":"f4e60b2e62b496c3714437bd100ab5236540abebd3685dfbd25eeddb37cb7160","kind":"language_definition","lines":"108-172","path":"include/Dataflow/IR/DataflowOps.td","roles":["context"],"text":"def Dataflow_StreamOp : Dataflow_Op<\"stream\", [Pure,\n CanonicalDataflowActor,\n AllTypesMatch<[\"init\", \"limit\", \"step\", \"iv\"]>]> {\n let summary = \"produce valid IV tokens and an explicit closing phase token\";\n let description = [{\n Executes one predicate-terminated integer recurrence per activation.\n `step_kind` selects the update applied after each true decision, and\n `predicate` compares the current value against `limit`.\n\n * `step_kind`: the canonical integer update operation: `add`, `sub`,\n `mul`, `sdiv`, `udiv`, `shl`, `ashr`, or `lshr`.\n * `predicate`: an upstream `arith.cmpi` predicate applied to the current\n value and `limit`.\n\n Idle consumes one `init` / `limit` / `step` triple. A true decision emits\n the current value on `iv`, emits `true` on `phase`, advances the current\n value, and remains active. A false decision emits only `false` on `phase`\n and returns to Idle. No sentinel IV is emitted, and subsequent activation\n triples may reuse the same operation.\n\n `init`, `limit`, `step`, and `iv` share one scalar signless integer type;\n `phase` is always `i1`. For `init=0`, `limit=5`, `step=1`, `step_kind=add`,\n and `predicate=slt`, `iv` is `0,1,2,3,4` and `phase` is `T,T,T,T,T,F`.\n }];\n\n let arguments = (ins\n AnySignlessInteger:$init,\n AnySignlessInteger:$limit,\n AnySignlessInteger:$step,\n Dataflow_StreamStepKindAttr:$step_kind,\n Arith_CmpIPredicateAttr:$predicate\n );\n let results = (outs AnySignlessInteger:$iv, I1:$phase);\n\n let hasCustomAssemblyFormat = 1;\n}\n\ndef Dataflow_CarryOp : Dataflow_Op<\"carry\", [Pure,\n CanonicalDataflowActor,\n AllTypesMatch<[\"init\", \"carry\", \"output\"]>]> {\n let summary =\n \"two-state carry: emit init once, then gate subsequent carries by cond\";\n let description = [{\n Two-state (init / carry) token-level element.\n\n * Start state is `init`.\n * In state `init`: wait for one `%init` token; forward it to\n `%output`; transition to `carry`.\n * In state `carry`: inspect the next `%cond : i1` token.\n - if `cond` is true : wait for and consume `%carry`, consume the\n condition, forward `%carry` to `%output`, and stay in `carry`;\n - if `cond` is false: consume only the condition, emit nothing, and\n return to state `init`.\n\n `%init`, `%carry` and `%output` share a single type; `%cond` is\n `i1`.\n }];\n\n let arguments = (ins I1:$cond, AnyType:$init, AnyType:$carry);\n let results = (outs AnyType:$output);\n\n let assemblyFormat = [{\n $cond `,` $init `,` $carry attr-dict `:` type($output)\n }];\n}","why":"dataflow.stream (iv plus closing phase token) and dataflow.carry (cond, init, carry -> output) operand order and semantics. The postcondition uses dataflow.stream to recognize a lowered loop and dataflow.carry as the recurrence element whose output must still gate each memory actor."},{"file_sha256":"f4e60b2e62b496c3714437bd100ab5236540abebd3685dfbd25eeddb37cb7160","kind":"language_definition","lines":"344-410","path":"include/Dataflow/IR/DataflowOps.td","roles":["context"],"text":"def Dataflow_SyncOp : Dataflow_Op<\"sync\", [Pure, CanonicalDataflowActor]> {\n let summary = \"wait for all inputs, then forward them all as outputs\";\n let description = [{\n Variadic rendezvous. Once every input has a token, all tokens are\n consumed and the corresponding outputs fire simultaneously.\n\n Operand count equals result count; types match positionally.\n }];\n\n let arguments = (ins Variadic:$inputs);\n let results = (outs Variadic:$outputs);\n\n let assemblyFormat = [{\n $inputs attr-dict `:` functional-type($inputs, $outputs)\n }];\n let hasVerifier = 1;\n}\n\ndef Dataflow_MuxOp : Dataflow_Op<\"mux\", [Pure, CanonicalDataflowActor]> {\n let summary = \"N-to-1 selection: forward the sel-picked input to output\";\n let description = [{\n N-input, 1-output multiplexer (N > 1). `sel` picks one input port\n index `k`. The op fires when tokens are available on *both* `%sel`\n and `%inputs[k]`. Both are consumed and the value is forwarded to\n `%output`. Tokens on the non-selected inputs are **not** consumed\n (they remain buffered; the other lanes block).\n\n `sel` type:\n * exactly 2 inputs -> `i1`\n * more than 2 inputs -> `index`\n\n All inputs and the output share one type.\n }];\n\n let arguments = (ins AnyTypeOf<[I1, Index]>:$sel,\n Variadic:$inputs);\n let results = (outs AnyType:$output);\n\n let assemblyFormat = [{\n $sel `,` $inputs attr-dict `:` functional-type(operands, results)\n }];\n let hasVerifier = 1;\n}\n\ndef Dataflow_DemuxOp : Dataflow_Op<\"demux\", [Pure, CanonicalDataflowActor]> {\n let summary = \"1-to-N selection: route the input to the sel-picked output\";\n let description = [{\n 1-input, N-output demultiplexer (N > 1). `sel` picks one output\n port index `k`. The op fires when tokens are available on both\n `%sel` and `%input`; both are consumed and the value is forwarded\n to `%outputs[k]`. No token is produced on the non-selected output\n ports.\n\n `sel` type:\n * exactly 2 outputs -> `i1`\n * more than 2 outputs -> `index`\n\n The input and all outputs share one type.\n }];\n\n let arguments = (ins AnyTypeOf<[I1, Index]>:$sel, AnyType:$input);\n let results = (outs Variadic:$outputs);\n\n let assemblyFormat = [{\n $sel `,` $input attr-dict `:` functional-type(operands, results)\n }];\n let hasVerifier = 1;","why":"dataflow.sync, dataflow.mux and dataflow.demux are the ordinary token-plumbing actors the frontier lanes pass through, so the postcondition's SSA token-flow closure must traverse them rather than assume a direct carry-to-actor edge."},{"file_sha256":"f4e60b2e62b496c3714437bd100ab5236540abebd3685dfbd25eeddb37cb7160","kind":"language_definition","lines":"412-524","path":"include/Dataflow/IR/DataflowOps.td","roles":["context"],"text":"//===----------------------------------------------------------------------===//\n// Memory Ops\n//\n// Streaming accesses against a memref, orchestrated by none-typed ctrl / done\n// tokens. The memref's element type constrains the element data type or the\n// access vector's element type.\n//===----------------------------------------------------------------------===//\n\n// The canonical memory actors. Each projects the standard MLIR memory effects\n// through the one shared implementation in `DataflowMemoryContracts.cpp`; no\n// actor classifies its own effects. That projection names the memory operand\n// for the addressed access and reads the atomic and volatile facts back from\n// the actor's one aggregate contract to add conservative unbound effects.\nclass Dataflow_MemoryActorOp traits = []>\n : Dataflow_Op])> {\n let extraClassDefinition = [{\n void $cppClass::getEffects(\n ::llvm::SmallVectorImpl<::mlir::MemoryEffects::EffectInstance>\n &effects) {\n ::dataflow::semantics::getMemoryActorEffects(getOperation(), effects);\n }\n }];\n}\n\ndef Dataflow_LoadOp : Dataflow_MemoryActorOp<\"load\"> {\n let summary = \"streaming element, contiguous vector, or gather load\";\n let description = [{\n On the simultaneous arrival of an address token and a `%ctrl : none`\n token, consumes both and fires one memory actor.\n A result type exactly equal to the memref element type loads one\n memory element, including when that element type is itself a vector.\n Otherwise, a fixed-size vector result of any positive rank loads that many\n elements in canonical row-major lane order, contiguously from a scalar\n `%addr : index` or, with a same-shape `%addr : vector<...xindex>`, one\n element per lane from the corresponding element-index address. A vector\n access requires the memref element type as its vector element type.\n\n An optional same-shape `i1` mask restricts a vector load to active lanes.\n Inactive lanes do not access memory and are deterministically zero-filled.\n After all active lanes retire, the op emits one data token and one `none`\n token on `%done`.\n\n The optional `contract` attribute is this actor's single\n `MemoryAccessContract`; its absence is the canonical plain non-volatile\n contract.\n }];\n\n let arguments = (ins AnyMemRef:$mem, AnyType:$addr, NoneType:$ctrl,\n Optional:$mask,\n OptionalAttr:$contract);\n let results = (outs AnyType:$data, NoneType:$done);\n\n let hasCustomAssemblyFormat = 1;\n let hasVerifier = 1;\n let builders = [\n OpBuilder<(ins\n \"::mlir::Type\":$data,\n \"::mlir::Type\":$done,\n \"::mlir::Value\":$mem,\n \"::mlir::Value\":$addr,\n \"::mlir::Value\":$ctrl)>,\n OpBuilder<(ins\n \"::mlir::Value\":$mem,\n \"::mlir::Value\":$addr,\n \"::mlir::Value\":$ctrl)>\n ];\n}\n\ndef Dataflow_StoreOp : Dataflow_MemoryActorOp<\"store\"> {\n let summary = \"streaming element, contiguous vector, or scatter store\";\n let description = [{\n On the simultaneous arrival of an address token, a `%data` value\n and a `%ctrl : none`, consumes all three and fires one memory actor.\n Data whose type exactly equals the memref element type writes one\n memory element, including when that element type is itself a vector.\n Otherwise, fixed-size vector data of any positive rank writes that many\n elements in canonical row-major lane order, contiguously from a scalar\n `%addr : index` or, with a same-shape `%addr : vector<...xindex>`, one\n element per lane. A vector access requires the memref element type as its\n vector element type.\n\n An optional same-shape `i1` mask restricts a vector store to active lanes.\n Inactive lanes do not access memory. After all active lanes retire, the op\n emits one `none` token on `%done`.\n\n The optional `contract` attribute is this actor's single\n `MemoryAccessContract`; its absence is the canonical plain non-volatile\n contract.\n }];\n\n let arguments = (ins AnyMemRef:$mem, AnyType:$addr, AnyType:$data,\n NoneType:$ctrl,\n Optional:$mask,\n OptionalAttr:$contract);\n let results = (outs NoneType:$done);\n\n let hasCustomAssemblyFormat = 1;\n let hasVerifier = 1;\n let builders = [\n OpBuilder<(ins\n \"::mlir::Type\":$done,\n \"::mlir::Value\":$mem,\n \"::mlir::Value\":$addr,\n \"::mlir::Value\":$data,\n \"::mlir::Value\":$ctrl)>,\n OpBuilder<(ins\n \"::mlir::Value\":$mem,\n \"::mlir::Value\":$addr,\n \"::mlir::Value\":$data,\n \"::mlir::Value\":$ctrl)>","why":"Canonical memory actors dataflow.load and dataflow.store: each consumes exactly one none-typed ctrl operand and produces a none-typed done result. This fixes how the postcondition identifies an actor's ctrl and done tokens by their none type."},{"file_sha256":"f4e60b2e62b496c3714437bd100ab5236540abebd3685dfbd25eeddb37cb7160","kind":"language_definition","lines":"839-877","path":"include/Dataflow/IR/DataflowOps.td","roles":["input_construction","input_well_formedness"],"text":"def Dataflow_GraphOp : Dataflow_Op<\"graph\", [\n IsolatedFromAbove,\n HasParent<\"::mlir::ModuleOp\">,\n SingleBlockImplicitTerminator<\"GraphReturnOp\">,\n FunctionOpInterface,\n RecursiveMemoryEffects,\n DeclareOpInterfaceMethods\n]> {\n let summary = \"Symbol-bearing function-like SpatialCore graph definition\";\n let description = [{\n Module-scope, function-like callable holding the SpatialCore body\n of a leaf dataflow graph. It does not itself execute; one or more\n `dataflow.graph.launch` ops materialise launches of it inside the\n body of a `dataflow.thread` definition.\n\n `function_type` contains only application payload ports. Normalized\n `input_segments` and `result_segments` classify those payloads as value,\n stream, and memory ports. The body's distinguished leading `none` block\n argument is the invocation start protocol endpoint, while launch `done`\n is derived exclusively from `dataflow.graph.return.complete`; neither is\n stored in the function type.\n\n This is the only canonical graph definition surface.\n }];\n\n let arguments = (ins\n SymbolNameAttr:$sym_name,\n TypeAttrOf:$function_type,\n DenseI32ArrayAttr:$input_segments,\n DenseI32ArrayAttr:$result_segments,\n OptionalAttr:$sym_visibility,\n OptionalAttr:$arg_attrs,\n OptionalAttr:$res_attrs);\n\n let regions = (region SizedRegion<1>:$body);\n\n let hasCustomAssemblyFormat = 1;\n let hasVerifier = 1;","why":"dataflow.graph is module-scope, single-block, symbol-bearing, carries required input_segments/result_segments payload classification and a distinguished leading none block argument that is not in the function type. The grammar spells every sampled graph exactly this way."},{"file_sha256":"f4e60b2e62b496c3714437bd100ab5236540abebd3685dfbd25eeddb37cb7160","kind":"language_definition","lines":"924-950","path":"include/Dataflow/IR/DataflowOps.td","roles":["input_construction","input_well_formedness"],"text":"def Dataflow_GraphReturnOp : Dataflow_Op<\"graph.return\", [\n AttrSizedOperandSegments,\n Terminator,\n ParentOneOf<[\"::dataflow::GraphOp\"]>,\n Pure\n]> {\n let summary = \"Terminator for a dataflow.graph body\";\n let description = [{\n Structurally declares the enclosing graph's value, stream, and memory\n outputs together with its mandatory retirement frontier. `complete` is\n an unordered all-of set of one or more `none` values; the launch `done`\n event is derived from that set and is not itself a return operand.\n\n The compact assembly form `%complete, %values... : none, types...` is\n retained for the common case with one completion witness and no stream\n or memory outputs. Other shapes print all four named segments.\n }];\n\n let arguments = (ins\n Variadic:$values,\n Variadic:$streams,\n Variadic:$memories,\n Variadic:$complete);\n\n let hasCustomAssemblyFormat = 1;\n\n let skipDefaultBuilders = 1;","why":"dataflow.graph.return segments (values, streams, memories, complete) and the retained compact form '%complete, %values... : none, types...' used by every sampled graph terminator."},{"file_sha256":"11d4b44ce36afb532b1ba720012841c38babad2962aadee232609a04fc28dbc4","kind":"test","lines":"1-27","path":"test/raise/scf-to-dfg-memory-frontier.mlir","roles":["input_construction"],"text":"// RUN: loom-raise-opt --split-input-file --loom-lower-graph-memory %s | FileCheck %s\n\n// CHECK-LABEL: dataflow.graph private @frontier_straight\n// CHECK: %[[R0:.*]], %[[D0:.*]] = dataflow.load %arg4[%arg1] %arg0 : memref<16xi32>\n// CHECK: %[[R1:.*]], %[[D1:.*]] = dataflow.load %arg4[%arg2] %arg0 : memref<16xi32>\n// CHECK: %[[WRITE:.*]] = dataflow.store %arg4[%arg1] %arg3 [[READS:%[^# ]+]]#0 : memref<16xi32>\n// CHECK: %[[R2:.*]], %[[D2:.*]] = dataflow.load %arg4[%arg2] %[[WRITE]] : memref<16xi32>\n// CHECK: [[READS]]:2 = dataflow.sync %[[D0]], %[[D1]] : (none, none) -> (none, none)\n// CHECK: %[[RB:.*]], %[[DB:.*]] = dataflow.load %arg5[%arg1] %[[WRITE]] : memref<16xi32>\n// CHECK: %[[RETIRE:.*]]:2 = dataflow.sync %[[D2]], %[[DB]] : (none, none) -> (none, none)\n// CHECK: dataflow.graph.return %[[RETIRE]]#0 : none\ndataflow.graph private @frontier_straight(\n %start: none, %i: index, %j: index, %value: i32,\n %a: memref<16xi32>, %b: memref<16xi32>) -> ()\n attributes {input_segments = array,\n result_segments = array} {\n %r0, %read0_done = dataflow.load %a[%i] %start : memref<16xi32>\n %r1, %read1_done = dataflow.load %a[%j] %start : memref<16xi32>\n %write_done = dataflow.store %a[%i] %value %start : memref<16xi32>\n %r2, %read2_done = dataflow.load %a[%j] %start : memref<16xi32>\n %rb = memref.load %b[%i] : memref<16xi32>\n dataflow.graph.return %start : none\n}\n\n// -----\n\n// Final values are published through the same explicit retirement frontier.","why":"Accepted input spelling for a finalized graph fed to --loom-lower-graph-memory: private visibility, leading %start: none, memref arguments, input_segments/result_segments arrays, and memref.load/store leaves in the body. The grammar mirrors this header shape."},{"file_sha256":"11d4b44ce36afb532b1ba720012841c38babad2962aadee232609a04fc28dbc4","kind":"test","lines":"106-170","path":"test/raise/scf-to-dfg-memory-frontier.mlir","roles":["input_construction","context"],"text":"// CHECK-LABEL: dataflow.graph private @frontier_for\n// CHECK: %[[IV:.*]], %[[PHASE:.*]] = dataflow.stream %arg1, %arg2, %arg3 step add while slt : i64\n// CHECK: %[[EXEC_RAW:.*]] = dataflow.carry %[[PHASE]], %arg0,\n// CHECK: %[[EXEC_LANES:.*]]:2 = dataflow.demux %[[PHASE]], %[[EXEC_RAW]] : (i1, none) -> (none, none)\n// CHECK: %[[VALUE_RAW:.*]] = dataflow.invariant %[[PHASE]], %arg5 : i32\n// CHECK: %[[BODY_PHASE:.*]], %[[BODY_VALUE:.*]] = dataflow.gate %[[PHASE]], %[[VALUE_RAW]] : i32\n// CHECK: %[[W_RAW:.*]] = dataflow.carry %[[PHASE]], %arg0,\n// CHECK: %[[R_RAW:.*]] = dataflow.carry %[[PHASE]], %arg0,\n// CHECK: %[[W_LANES:.*]]:2 = dataflow.demux %[[PHASE]], %[[W_RAW]] : (i1, none) -> (none, none)\n// CHECK: %[[R_LANES:.*]]:2 = dataflow.demux %[[PHASE]], %[[R_RAW]] : (i1, none) -> (none, none)\n// CHECK: dataflow.load %arg6[{{.*}}]\n// CHECK: %[[STORE_DONE:.*]] = dataflow.store %arg6[{{.*}}] %[[BODY_VALUE]]\n// CHECK: dataflow.load %arg6[%arg4]\n// CHECK-NOT: scf.for\ndataflow.graph private @frontier_for(\n %start: none, %lb: i64, %ub: i64, %step: i64,\n %after_index: index, %value: i32,\n %a: memref, %b: memref) -> ()\n attributes {input_segments = array,\n result_segments = array} {\n scf.for %i = %lb to %ub step %step : i64 {\n %index = arith.index_cast %i : i64 to index\n %loaded = memref.load %a[%index] : memref\n memref.store %value, %a[%index] : memref\n }\n %after = memref.load %a[%after_index] : memref\n dataflow.graph.return %start : none\n}\n\n// CHECK-LABEL: dataflow.graph private @frontier_for_zero_trip\n// CHECK: %[[ZERO_IV:.*]], %[[ZERO_PHASE:.*]] = dataflow.stream %arg1, %arg1, %arg2 step add while slt : i64\n// CHECK: %[[ZERO_EXEC_RAW:.*]] = dataflow.carry %[[ZERO_PHASE]], %arg0,\n// CHECK: %[[ZERO_EXEC_LANES:.*]]:2 = dataflow.demux %[[ZERO_PHASE]], %[[ZERO_EXEC_RAW]] : (i1, none) -> (none, none)\n// CHECK: %[[ZERO_VALUE_RAW:.*]] = dataflow.carry %[[ZERO_PHASE]], %arg4,\n// CHECK: %[[ZERO_VALUE_LANES:.*]]:2 = dataflow.demux %[[ZERO_PHASE]], %[[ZERO_VALUE_RAW]] : (i1, i32) -> (i32, i32)\n// CHECK: %[[ZERO_INDEX_RAW:.*]] = dataflow.invariant %[[ZERO_PHASE]], %arg3 : index\n// CHECK: %[[ZERO_BODY_PHASE:.*]], %[[ZERO_BODY_VALUE:.*]] = dataflow.gate %[[ZERO_PHASE]], %[[ZERO_INDEX_RAW]] : index\n// CHECK: %[[ZERO_BODY_CLOSE:.*]]:2 = dataflow.demux %[[ZERO_BODY_PHASE]], %[[ZERO_BODY_VALUE]] : (i1, index) -> (index, index)\n// CHECK: %[[ZERO_W_RAW:.*]] = dataflow.carry %[[ZERO_PHASE]], %arg0,\n// CHECK: %[[ZERO_R_RAW:.*]] = dataflow.carry %[[ZERO_PHASE]], %arg0,\n// CHECK: %[[ZERO_W_LANES:.*]]:2 = dataflow.demux %[[ZERO_PHASE]], %[[ZERO_W_RAW]] : (i1, none) -> (none, none)\n// CHECK: %[[ZERO_R_LANES:.*]]:2 = dataflow.demux %[[ZERO_PHASE]], %[[ZERO_R_RAW]] : (i1, none) -> (none, none)\n// CHECK: %[[ZERO_NONEMPTY:.*]] = arith.cmpi slt, %arg1, %arg1\n// CHECK: %[[ZERO_COMPLETION_LANES:.*]]:2 = dataflow.demux %[[ZERO_NONEMPTY]], %[[ZERO_EXEC_LANES]]#0 : (i1, none) -> (none, none)\n// CHECK: %[[ZERO_ACTIVE_RETIRE:.*]]:2 = dataflow.sync %[[ZERO_COMPLETION_LANES]]#1, %[[ZERO_BODY_CLOSE]]#0 : (none, index) -> (none, index)\n// CHECK: %[[ZERO_EXEC_RETIRE:.*]] = dataflow.mux %[[ZERO_NONEMPTY]], %[[ZERO_COMPLETION_LANES]]#0, %[[ZERO_ACTIVE_RETIRE]]#0 : (i1, none, none) -> none\n// CHECK: %[[ZERO_AFTER_CTRL:.*]]:2 = dataflow.sync %[[ZERO_EXEC_RETIRE]], %[[ZERO_W_LANES]]#0 : (none, none) -> (none, none)\n// CHECK: %{{.*}}, %[[ZERO_LOAD_DONE:.*]] = dataflow.load %arg5[%arg3] %[[ZERO_AFTER_CTRL]]#0 : memref\n// CHECK: %[[ZERO_MEMORY_RETIRE:.*]]:2 = dataflow.sync %[[ZERO_R_LANES]]#0, %[[ZERO_LOAD_DONE]] : (none, none) -> (none, none)\n// CHECK: %[[ZERO_RETIRE:.*]]:2 = dataflow.sync %[[ZERO_MEMORY_RETIRE]]#0, %[[ZERO_VALUE_LANES]]#0 : (none, i32) -> (none, i32)\n// CHECK: dataflow.graph.return %[[ZERO_RETIRE]]#0, %[[ZERO_RETIRE]]#1 : none, i32\ndataflow.graph private @frontier_for_zero_trip(\n %start: none, %bound: i64, %step: i64, %index: index, %value: i32,\n %a: memref) -> (i32)\n attributes {input_segments = array,\n result_segments = array} {\n %result = scf.for %i = %bound to %bound step %step\n iter_args(%state = %value) -> (i32) : i64 {\n memref.store %state, %a[%index] : memref\n scf.yield %state : i32\n }\n %after = memref.load %a[%index] : memref\n dataflow.graph.return %start, %result : none, i32\n}","why":"The frontier_for and frontier_for_zero_trip cases show an accepted source-sequential scf.for input (i64 induction, index_cast, memref accesses, iter_args) and the expected lowered shape in which dataflow.carry rings for W and R take the body's memory done as feedback while the body actors take ctrl from the ring lanes. Used to construct the sampled loop bodies and to confirm the ring terminology the postcondition tests."},{"file_sha256":"11d4b44ce36afb532b1ba720012841c38babad2962aadee232609a04fc28dbc4","kind":"test","lines":"317-343","path":"test/raise/scf-to-dfg-memory-frontier.mlir","roles":["input_construction"],"text":"// CHECK-LABEL: dataflow.graph private @frontier_nested_for_while\n// CHECK: dataflow.stream\n// CHECK: dataflow.carry\n// CHECK: dataflow.carry\n// CHECK: dataflow.load\n// CHECK-NOT: scf.for\n// CHECK-NOT: scf.while\ndataflow.graph private @frontier_nested_for_while(\n %start: none, %lb: i64, %ub: i64, %step: i64,\n %limit: i64, %one: i64, %a: memref) -> ()\n attributes {input_segments = array,\n result_segments = array} {\n scf.for %outer = %lb to %ub step %step : i64 {\n %result = scf.while (%inner = %outer) : (i64) -> i64 {\n %index = arith.index_cast %inner : i64 to index\n %loaded = memref.load %a[%index] : memref\n %continue = arith.cmpi slt, %inner, %limit : i64\n scf.condition(%continue) %inner : i64\n } do {\n ^bb0(%after: i64):\n %next = arith.addi %after, %one : i64\n scf.yield %next : i64\n }\n }\n dataflow.graph.return %start : none\n}","why":"Accepted nested source-sequential loop input (scf.for containing scf.while / scf.for), evidence that the nested loop shape sampled by the grammar is inside the supported domain."}],"primary_bundle_sha256":"ab5a589d3845298be90a59961e11f9ec1fabbd12e7629b74a5b9262e796e72fe","project":"PolyArch/loom","revision":"48615bc5925ef4b9db8b4550b5d4322933cf4b7b","schema":"spectriad.authoring-context/v1","selection_sha256":"323bfca6272ef97e1f0a2fba92220a91050cfda088269cf715ef488f9dae0dcf"}