{"entries":[{"file_sha256":"d76b4cb1e888697d5f011a939e68cbc6230647c4d457c12d22689740aa43a44d","kind":"documentation_input","lines":"988-1037","path":"docs/spec-compiler-part-3-dfg.md","roles":["context","input_construction","input_well_formedness"],"text":"* `dataflow.thread` is a Symbol-bearing, module-scope, function-\n like callable. It does not itself execute; one or more\n `dataflow.thread.launch` ops materialize launches of it.\n* `function_type` is a `FunctionType` whose inputs are the kernel's\n user-data operand types `(T0, ..., TN)` and whose results are\n empty. The thread definition has no SSA data results; the\n per-launch completion token is launch-side, not part of the\n callable signature. Asynchronous execution is expressed by launch\n dependencies and the mandatory launch completion token, not by the\n function type.\n* `sym_name` is required and module-unique. `sym_visibility` is\n required and must equal `\"private\"` under the baseline visibility\n policy. The verifier rejects `\"public\"` and `\"nested\"` unless\n cross-module linkage is enabled by a separate spec.\n* `domain` is the closed `DenseRectangular` or\n `DynamicWork { work_item_arg_ordinal }` value owned by Part 4. A dynamic\n definition has coordinate rank zero and its ordinal must select exactly one\n `function_type` input. A dense definition has no work-item ordinal.\n* A dense entry block has the layout `(args_*, thread_ctrl, coord_*)`; a\n dynamic entry block has `(args_*, thread_ctrl)`:\n - The first `N` block arguments mirror `function_type.inputs`\n exactly (each user body operand). Putting the signature args\n first preserves the upstream `FunctionOpInterface` invariant\n that the entry block's first `N` arguments correspond to\n `function_type.inputs[0..N]`. This matches the `gpu.func`\n precedent of \"function args first, implicit extras after\".\n - `thread_ctrl : none` is the per-launch AccCore start signal.\n It is produced by the launch op once async dependencies are\n satisfied and the AccCore instance begins execution. Root\n `dataflow.graph.launch` ops with no InstructionCore predecessor use\n this value as their `ctrl_in` operand.\n - For a dense definition, `coord_0, ..., coord_{K-1} : index` are the\n per-instance logical\n coordinates, one per launch-domain dimension, in source-dimension order.\n Their count is the definition's coordinate rank. Rank is derived from\n this canonical suffix after the `function_type` inputs and unique\n `thread_ctrl`; there is no duplicate rank, grid, or mapping attribute.\n - For a dynamic definition, the designated ordinary argument carries the\n current work-item payload. The runtime `WorkItemId` is execution identity,\n not an additional SSA argument or payload wrapper.\n* `arg_attrs` is indexed only by the `args_*` payload prefix. Forall\n promotion copies each captured source function argument dictionary,\n including arbitrary attributes such as `llvm.noalias`, into capture order.\n Locally defined captures have an empty dictionary. `thread_ctrl` and\n `coord_*` are not payload arguments and never inherit source argument\n metadata.\n* The body is `IsolatedFromAbove`. No SSA value defined outside\n the def's body may be used inside it; the launch's body operands\n are the only inputs.\n#### 5.4.2 `dataflow.thread.launch`","why":"Governing context for dataflow.thread: private visibility, result-free function_type, closed dense/dynamic domain, and the (args_*, thread_ctrl, coord_*) entry-block layout that the generated definition-and-launch carrier inputs must instantiate."},{"file_sha256":"d76b4cb1e888697d5f011a939e68cbc6230647c4d457c12d22689740aa43a44d","kind":"documentation_input","lines":"2008-2023","path":"docs/spec-compiler-part-3-dfg.md","roles":["input_construction","input_well_formedness"],"text":"* retained a graph-owned forall only as a mapping-free, effect-form,\n compile-time fixed-domain construct whose `P[]` width, ownership, and\n cross-lane legality are materialized in semantic IR and can be re-proved;\n and\n* materialized every supported aggregation or reduction into accepted\n semantics, or failed finalizability truthfully.\n\nPart 3 does not convert aggregation form, decide thread ownership, rewrite\nforall to parallel as an optimization policy, infer `P[]`, serialize lanes, or\nselect a reduction strategy. A dynamic domain, mapping attribute, shared\noutput, result, combining action, or failed legality re-proof causes atomic\nfailure before canonical graph publication. Cached provenance never changes\nthis result.\n\nFor this boundary, an accepted effect-form forall has no `shared_outs`, no op\nresults, and an empty `scf.forall.in_parallel` terminator. In the example,","why":"Fixes the accepted forall form carried into Part 3: mapping-free, effect-form, compile-time fixed domain, no shared_outs, no results, empty scf.forall.in_parallel terminator, and no dynamic domain."},{"file_sha256":"d76b4cb1e888697d5f011a939e68cbc6230647c4d457c12d22689740aa43a44d","kind":"documentation_input","lines":"2052-2061","path":"docs/spec-compiler-part-3-dfg.md","roles":["input_construction","applicability"],"text":"Part 3 rejects this form before graph mutation. It never drops the combining\nregion or publishes a `dataflow.graph` that omits the aggregation. Any legal\nmaterialization belongs to the Part 2 owner; this document intentionally does\nnot define a bufferization or combining algorithm.\n\nIf Part 2 selects an effect-form forall as an AccCore thread domain, the input\naccepted by Part 3 is already the definition-and-launch carrier shape below.\nThe rank-one source sketch is retained only to relate the source induction\nvariable to the canonical logical-coordinate ABI; it is not a Part 3\ntransformation:","why":"States that a forall selected as an AccCore thread domain reaches Part 3 already in the definition-and-launch carrier shape, which is the input shape the grammar samples."},{"file_sha256":"d76b4cb1e888697d5f011a939e68cbc6230647c4d457c12d22689740aa43a44d","kind":"documentation_input","lines":"2096-2100","path":"docs/spec-compiler-part-3-dfg.md","roles":["input_construction"],"text":"Code inside the thread definition remains InstructionCore code unless the\nselected Structured Program Candidate explicitly wraps it in\n`loom.spatial_region`. That compiler-internal region remains the temporary\nSpatialCore ownership carrier until Part 3 atomically replaces it with a\nfinalized `dataflow.graph` definition and launch. Memory operations outside","why":"Code inside a thread definition stays InstructionCore code unless explicitly wrapped in loom.spatial_region, justifying thread bodies that hold plain scf/memref code without a spatial region."},{"file_sha256":"bfc1e646e91fe0ba6d7d16e43994100b05c3b8ff79c2e07288e8955c43d9d79d","kind":"documentation_input","lines":"934-942","path":"docs/spec-compiler-part-2-scf.md","roles":["input_construction","input_well_formedness"],"text":"temporary joint proof do not survive in the child. When the selected nest is\ninside an already materialized rank-zero Spatial ownership carrier, the same\natomic decision promotes the new forall to the carrier's dense logical thread\ndomain. Ownership remains the sole owner of extent arithmetic and\nsource-induction reconstruction: every exact thread launch must project all\nbounds from its body operands, the thread and Spatial ABIs acquire the rank-N\ncoordinate suffix, and no graph-owned forall remains. An unprojectable bound\nrejects only that candidate. The transformation never weakens the fixed-domain\nrequirement for a retained graph-owned parallel form.","why":"Promotion into a dense logical thread domain and the fixed-domain requirement for a retained graph-owned parallel form; drives the zero-based constant forall extents and dense domain attribute in generated inputs."},{"file_sha256":"f4e60b2e62b496c3714437bd100ab5236540abebd3685dfbd25eeddb37cb7160","kind":"language_definition","lines":"614-739","path":"include/Dataflow/IR/DataflowOps.td","roles":["input_construction","input_well_formedness"],"text":"def Dataflow_ThreadOp : Dataflow_Op<\"thread\", [\n AutomaticAllocationScope,\n IsolatedFromAbove,\n HasParent<\"::mlir::ModuleOp\">,\n SingleBlockImplicitTerminator<\"ThreadYieldOp\">,\n FunctionOpInterface,\n RecursiveMemoryEffects\n]> {\n let summary = \"Symbol-bearing function-like AccCore kernel definition\";\n let description = [{\n Module-scope, function-like callable that holds an AccCore kernel\n body. It does not itself execute; one or more\n `dataflow.thread.launch` ops materialize launches of it.\n\n The body's entry block has the layout\n `(args_*, thread_ctrl: none, iv_*: index)` (per spec section\n 5.4.1). The first N block args mirror `function_type.inputs`; the\n trailing `thread_ctrl` and grid index args are NOT in\n `function_type` (they are launch-instance extras). Specifically:\n\n * Args[0 .. N-1] match `function_type.inputs` position-wise.\n * Args[N] is `none` -- the per-launch `thread_ctrl`\n slot, used as the AccCore start signal and\n consumed by root `dataflow.graph.launch` ops\n in the body as a dependency event.\n * Args[N+1 .. end] are all `index` -- one per grid dim.\n\n The custom assembly format prints the required `domain(...)` immediately\n after the symbol and the trailing extras after the function-style\n signature using a separate `ctrl ( ... )` clause\n (the `thread_ctrl` slot) and an `iv ( ... )` clause (the\n grid-index slots). Either / both clauses are optional; threads\n written without them are accepted at parse time only when the op\n is external (i.e., body is empty), since a body-having thread\n must carry the trailing `thread_ctrl` slot per the verifier.\n\n The op is `IsolatedFromAbove`; values flow in only through the\n matching `dataflow.thread.launch` body operands.\n\n Every definition carries one closed `domain`: DenseRectangular or\n DynamicWork. Dense rank is derived solely from the trailing index block\n arguments. DynamicWork carries one ordinary function-input ordinal and has\n no coordinate suffix.\n }];\n\n let arguments = (ins\n SymbolNameAttr:$sym_name,\n TypeAttrOf:$function_type,\n Dataflow_ThreadDomainAttr:$domain,\n OptionalAttr:$sym_visibility,\n OptionalAttr:$arg_attrs,\n OptionalAttr:$res_attrs);\n\n let regions = (region SizedRegion<1>:$body);\n\n let hasCustomAssemblyFormat = 1;\n let hasVerifier = 1;\n\n let builders = [\n OpBuilder<(ins\n \"::llvm::StringRef\":$name,\n \"::mlir::FunctionType\":$type,\n \"::dataflow::ThreadDomainAttr\":$domain,\n CArg<\"::llvm::ArrayRef<::mlir::NamedAttribute>\", \"{}\">:$attrs)>\n ];\n\n let extraClassDeclaration = [{\n /// FunctionOpInterface methods.\n ::llvm::ArrayRef<::mlir::Type> getArgumentTypes() {\n return getFunctionType().getInputs();\n }\n ::llvm::ArrayRef<::mlir::Type> getResultTypes() {\n return getFunctionType().getResults();\n }\n ::mlir::Region *getCallableRegion() {\n return isExternal() ? nullptr : &getBody();\n }\n bool isExternal() { return getBody().empty(); }\n\n /// Override the default FunctionOpInterface body verifier: a\n /// dataflow.thread body's entry block leads with the\n /// function-signature args, then a `none`-typed `thread_ctrl`\n /// slot, then zero or more `index`-typed `iv_*` slots (per spec\n /// section 5.4.1: `(args_*, thread_ctrl, iv_*)`).\n ::llvm::LogicalResult verifyBody() {\n if (isExternal())\n return ::mlir::success();\n ::mlir::Block &entry = getBody().front();\n ::llvm::ArrayRef<::mlir::Type> inputs = getFunctionType().getInputs();\n const size_t N = inputs.size();\n // Body-carrying threads MUST have the trailing thread_ctrl slot.\n if (entry.getNumArguments() < N + 1)\n return emitOpError(\"entry block must have at least \")\n << (N + 1)\n << \" arguments (function inputs + 1 thread_ctrl slot)\";\n // First N entry block arguments must match function_type.inputs.\n for (size_t i = 0; i < N; ++i) {\n if (entry.getArgument(i).getType() != inputs[i])\n return emitOpError(\"type of entry block argument #\")\n << i << '(' << entry.getArgument(i).getType()\n << \") must match the corresponding function signature input (\"\n << inputs[i] << ')';\n }\n // Slot N must be the `none`-typed thread_ctrl per spec\n // section 5.4.1.\n if (!::llvm::isa<::mlir::NoneType>(entry.getArgument(N).getType()))\n return emitOpError(\"entry block argument #\")\n << N << \" (thread_ctrl) must have type `none`, got \"\n << entry.getArgument(N).getType();\n if (getDomain().getKind() ==\n ::dataflow::ThreadDomainKind::DynamicWork &&\n entry.getNumArguments() != N + 1)\n return emitOpError(\"dynamic-work thread body must not have coordinate \"\n \"arguments\");\n // All remaining dense-domain slots are grid induction variables of\n // `index`.\n for (size_t i = N + 1, e = entry.getNumArguments(); i < e; ++i) {\n if (!::llvm::isa<::mlir::IndexType>(entry.getArgument(i).getType()))\n return emitOpError(\"entry block argument #\")\n << i << \" (grid iv) must have type `index`, got \"\n << entry.getArgument(i).getType();\n }\n return ::mlir::success();\n }\n }];\n}","why":"Dataflow_ThreadOp definition: operands/attributes (sym_name, function_type, domain, sym_visibility), the custom `domain(...) (...) ctrl (...) iv (...)` assembly clauses, and verifyBody's entry-block prefix/ctrl/index slot requirements that generated inputs must satisfy."},{"file_sha256":"f4e60b2e62b496c3714437bd100ab5236540abebd3685dfbd25eeddb37cb7160","kind":"language_definition","lines":"741-820","path":"include/Dataflow/IR/DataflowOps.td","roles":["input_construction","input_well_formedness"],"text":"def Dataflow_ThreadYieldOp : Dataflow_Op<\"thread.yield\", [\n Terminator,\n ParentOneOf<[\"::dataflow::ThreadOp\"]>,\n Pure\n]> {\n let summary = \"Terminator for a dataflow.thread body\";\n let description = [{\n Accepts an unordered all-of completion frontier of `none` values.\n Tensor-result aggregation from `scf.forall` is materialised into\n explicit destination-buffer writes before thread promotion, so the\n thread definition has no parallel combining region or thread data\n results.\n }];\n\n let arguments = (ins Variadic:$completionFrontier);\n\n let assemblyFormat = \"($completionFrontier^ `:` type($completionFrontier))? attr-dict\";\n\n let skipDefaultBuilders = 1;\n let builders = [\n OpBuilder<(ins CArg<\"::mlir::ValueRange\", \"{}\">:$completionFrontier), [{\n $_state.addOperands(completionFrontier);\n }]>\n ];\n}\n\ndef Dataflow_ThreadLaunchOp : Dataflow_Op<\"thread.launch\", [\n AttrSizedOperandSegments,\n DeclareOpInterfaceMethods\n]> {\n let summary = \"Async launch of a dataflow.thread callable\";\n let description = [{\n References a `dataflow.thread` definition by symbol and supplies\n body operands and optional grid upper bounds. Always produces one\n `!dataflow.thread_token` completion handle for all dynamic launch\n instances.\n\n Grid lower bounds and steps are not modeled; body operands, upper bounds,\n and async dependencies are explicit.\n }];\n\n let arguments = (ins\n FlatSymbolRefAttr:$callee,\n Variadic:$bodyOperands,\n Variadic:$gridUpperBounds,\n Variadic:$asyncDependencies);\n\n let results = (outs Dataflow_ThreadTokenType:$asyncToken);\n\n let assemblyFormat = [{\n $callee\n `(` $bodyOperands `)`\n ( `grid` `(` $gridUpperBounds^ `)` )?\n ( `wait` `(` $asyncDependencies^ `)` )?\n attr-dict `:`\n functional-type($bodyOperands, results)\n }];\n\n let skipDefaultBuilders = 1;\n let builders = [\n OpBuilder<(ins\n \"::mlir::FlatSymbolRefAttr\":$callee,\n \"::mlir::ValueRange\":$bodyOperands,\n \"::mlir::ValueRange\":$gridUpperBounds,\n \"::mlir::ValueRange\":$asyncDependencies), [{\n $_state.addOperands(bodyOperands);\n $_state.addOperands(gridUpperBounds);\n $_state.addOperands(asyncDependencies);\n auto &properties = $_state.getOrAddProperties();\n properties.operandSegmentSizes = {\n static_cast(bodyOperands.size()),\n static_cast(gridUpperBounds.size()),\n static_cast(asyncDependencies.size())};\n properties.callee = callee;\n $_state.addTypes(::dataflow::ThreadTokenType::get($_builder.getContext()));\n }]>\n ];\n\n let hasVerifier = 1;\n}","why":"Dataflow_ThreadYieldOp terminator and Dataflow_ThreadLaunchOp assembly format (callee, body operands, optional grid and wait clauses, thread_token result) used to spell the launch side of each sampled carrier."},{"file_sha256":"0616db64bbc547b2c92dd9801dbf1dac5136f19fc11ebdd1019ffd9660af9534","kind":"verifier","lines":"413,563-584,602-639","path":"lib/Dataflow/IR/DataflowFunctionLikeOps.cpp","roles":["input_well_formedness"],"text":"ParseResult ThreadOp::parse(OpAsmParser &parser, OperationState &result) {\nLogicalResult ThreadOp::verify() {\n if (!getSymVisibility() || *getSymVisibility() != \"private\")\n return emitOpError(\"requires explicit 'private' visibility\");\n if (getFunctionType().getNumResults() != 0)\n return emitOpError(\"must not declare function results\");\n\n if (getDomain().getKind() == ThreadDomainKind::DynamicWork) {\n uint64_t ordinal = *getDomain().getWorkItemArgOrdinal();\n ArrayRef inputs = getFunctionType().getInputs();\n if (ordinal >= inputs.size())\n return emitOpError(\"dynamic-work item argument ordinal \")\n << ordinal << \" is out of bounds for \" << inputs.size()\n << \" thread inputs\";\n for (auto [index, type] : llvm::enumerate(inputs))\n if (DataflowDialect::containsChannelOrThreadToken(type))\n return emitOpError(\"dynamic-work thread input #\")\n << index << \" must not contain a channel or thread token\";\n }\n if (!ownsThreadLaunchExtentAnalysis(*this))\n return success();\n return verifyThreadLaunchExtents(cast((*this)->getParentOp()));\n}\nLogicalResult ThreadLaunchOp::verifySymbolUses(SymbolTableCollection &symbols) {\n auto callee =\n symbols.lookupNearestSymbolFrom(*this, getCalleeAttr());\n if (!callee)\n return emitOpError(\"'\")\n << getCallee()\n << \"' does not reference a valid 'dataflow.thread' op\";\n\n // Body operand types must equal callee.function_type.inputs\n // position-by-position.\n ArrayRef calleeInputs = callee.getFunctionType().getInputs();\n if (getBodyOperands().size() != calleeInputs.size())\n return emitOpError(\"body operand count (\")\n << getBodyOperands().size()\n << \") does not match callee input count (\" << calleeInputs.size()\n << \")\";\n for (size_t i = 0, e = calleeInputs.size(); i < e; ++i) {\n Type expected = calleeInputs[i];\n Type actual = getBodyOperands()[i].getType();\n if (actual != expected)\n return emitOpError(\"body operand #\")\n << i << \" type \" << actual << \" does not match callee input type \"\n << expected;\n }\n\n size_t calleeRank = 0;\n if (!callee.isExternal()) {\n size_t entryArgumentCount = callee.getBody().front().getNumArguments();\n size_t requiredArgumentCount = calleeInputs.size() + 1;\n if (entryArgumentCount >= requiredArgumentCount)\n calleeRank = entryArgumentCount - requiredArgumentCount;\n }\n if (getGridUpperBounds().size() != calleeRank)\n return emitOpError(\"grid upper bound count (\")\n << getGridUpperBounds().size() << \") must match callee rank (\"\n << calleeRank << \")\";\n return success();\n}","why":"ThreadOp::parse derives function_type from the printed signature; ThreadOp::verify requires explicit private visibility and no function results; ThreadLaunchOp::verifySymbolUses requires body-operand types to equal callee inputs and grid count to equal callee coordinate rank. These bound which generated modules the subject accepts."},{"file_sha256":"0616db64bbc547b2c92dd9801dbf1dac5136f19fc11ebdd1019ffd9660af9534","kind":"verifier","lines":"313-347","path":"lib/Dataflow/IR/DataflowFunctionLikeOps.cpp","roles":["input_well_formedness"],"text":"static LogicalResult verifyThreadLaunchExtents(ModuleOp module) {\n ExtentConstantEvaluator evaluator;\n WalkResult result =\n module.walk([&](Operation *op) -> WalkResult {\n if (op != module.getOperation() && op->hasTrait())\n return WalkResult::skip();\n\n auto launch = dyn_cast(op);\n if (!launch)\n return WalkResult::advance();\n // A launch whose segmentation cannot be read safely is diagnosed by\n // its own verification; this analysis just leaves it alone.\n if (!hasReadableOperandSegments(launch))\n return WalkResult::advance();\n for (auto [index, extent] :\n llvm::enumerate(launch.getGridUpperBounds())) {\n auto constant =\n dyn_cast_or_null(evaluator.evaluate(extent));\n if (constant && constant.getValue().isNegative()) {\n launch.emitOpError(\"grid upper bound #\")\n << index << \" must be nonnegative\";\n return WalkResult::interrupt();\n }\n }\n return WalkResult::advance();\n });\n return success(!result.wasInterrupted());\n}\n\nstatic bool ownsThreadLaunchExtentAnalysis(ThreadOp thread) {\n for (Operation *previous = thread->getPrevNode(); previous;\n previous = previous->getPrevNode())\n if (isa(previous))\n return false;\n return true;","why":"Module-scope thread launch extent analysis rejecting negative constant grid upper bounds; generated grid bounds are nonnegative index constants."},{"file_sha256":"c6f79ffa52853dc0bfac5214fbc7f94b0dacc45f8e5cdf216b2f0812c02bdff2","kind":"implementation","lines":"1-70","path":"lib/Frontend/Lowering/LowerForallToThreadPass.cpp","roles":["applicability","context"],"text":"// Reject implicit thread ownership for host scf.forall operations.\n//\n// dataflow.thread.launch models only zero-based grid extents. Thread\n// promotion therefore requires a recognized Loom mapping and a prior\n// structured-domain transformation that preserves lower bounds and steps.\n// Neither authority is currently materialized by this pass.\n\n#include \"Frontend/Lowering/Passes.h\"\n\n#include \"mlir/Dialect/Func/IR/FuncOps.h\"\n#include \"mlir/Dialect/SCF/IR/SCF.h\"\n#include \"mlir/IR/BuiltinOps.h\"\n#include \"mlir/Pass/Pass.h\"\n#include \"mlir/Pass/PassRegistry.h\"\n\nnamespace {\n\nstruct LowerForallToThreadPass\n : public ::mlir::PassWrapper> {\n MLIR_DEFINE_EXPLICIT_INTERNAL_INLINE_TYPE_ID(LowerForallToThreadPass)\n\n ::llvm::StringRef getArgument() const final {\n return \"loom-lower-forall-to-thread\";\n }\n\n ::llvm::StringRef getDescription() const final {\n return \"Reject scf.forall thread promotion without a recognized Loom \"\n \"mapping and faithfully represented domain.\";\n }\n\n void getDependentDialects(::mlir::DialectRegistry ®istry) const final {\n registry.insert<::mlir::func::FuncDialect, ::mlir::scf::SCFDialect>();\n }\n\n void runOnOperation() final {\n bool rejected = false;\n getOperation().walk([&](::mlir::scf::ForallOp forall) {\n if (!forall->getParentOfType<::mlir::func::FuncOp>())\n return;\n forall.emitError(\n \"loom-lower-forall-to-thread: raw scf.forall has no recognized \"\n \"Loom thread mapping; preserve it until structured ownership and \"\n \"its complete domain are selected\");\n rejected = true;\n });\n if (rejected)\n signalPassFailure();\n }\n};\n\n} // namespace\n\nnamespace loom {\nnamespace lowering {\n\nstd::unique_ptr<::mlir::Pass> createLowerForallToThreadPass() {\n return std::make_unique();\n}\n\nvoid registerLowerForallToThreadPass() {\n static bool once = []() {\n ::mlir::PassRegistration();\n return true;\n }();\n (void)once;\n}\n\n} // namespace lowering\n} // namespace loom","why":"The selected stage only diagnoses scf.forall ops that have a func.func ancestor and never materializes a dataflow.thread; this fixes which sampled inputs the stage accepts (forall inside the thread carrier) and is the evidence for the reported stage-attribution mismatch."},{"file_sha256":"7f380008fb405f6cf8d60d16d1b5980dc1d2e694f9842bcfeb6538fbcfa01099","kind":"implementation","lines":"1-30","path":"lib/Frontend/Lowering/Pipeline.cpp","roles":["applicability","context"],"text":"// Pipeline glue and pass-registry hooks for the SCF-to-DFG lowering\n// passes. The standard pipeline runs:\n//\n// loom-lower-for-to-graph (module-level)\n//\n// `loom-lower-for-to-graph` owns the atomic publication transaction. It\n// consumes explicit loom.spatial_region candidates, runs graph finalization\n// on a scratch module, validates the native result, and publishes only the\n// completed module. Graph memref-copy expansion is part of that finalization,\n// so a copy the current profile cannot expand fails the transaction instead of\n// reaching the published program.\n//\n// Thread ownership must already be present in the Structured Program\n// Candidate. The independently registered forall pass only diagnoses raw\n// implicit promotion requests.\n\n#include \"Frontend/Lowering/Passes.h\"\n\n#include \"mlir/Pass/PassManager.h\"\n#include \"mlir/Pass/PassRegistry.h\"\n#include \"mlir/Transforms/Passes.h\"\n\nnamespace loom {\nnamespace lowering {\n\nvoid registerExpandGraphMemrefCopyPass();\nvoid registerLowerForallToThreadPass();\nvoid registerLowerForToGraphPass();\nvoid registerLowerGraphConstantsPass();\nvoid registerLowerGraphMemoryPass();","why":"Documents that thread ownership must already be present in the Structured Program Candidate and that the forall pass only diagnoses raw implicit promotion requests; supports keeping the stage flag and sampling pre-formed carriers."},{"file_sha256":"634d4f13677fbd71bdd190bbb48feb1619c6d88f387330c467c4e1fe84203407","kind":"test","lines":"1-26","path":"test/raise/scf-to-dfg-forall-to-thread.mlir","roles":["applicability"],"text":"// RUN: not loom-raise-opt --loom-lower-forall-to-thread \\\n// RUN: --mlir-disable-threading --mlir-print-ir-after-failure \\\n// RUN: --mlir-print-ir-module-scope %s 2>&1 | FileCheck %s\n\n// Thread promotion must not infer ownership for an unmapped forall. The\n// thread launch ABI cannot represent this offset, strided domain, so failure\n// must preserve the complete source domain rather than treating the upper\n// bound as a zero-based grid extent.\n\n// CHECK: error: loom-lower-forall-to-thread: raw scf.forall has no recognized Loom thread mapping\n// CHECK-LABEL: func.func @offset_strided(\n// CHECK: scf.forall\n// CHECK-SAME: (%{{.*}}) = (%{{.*}}) to (%{{.*}}) step (%{{.*}})\n// CHECK-NOT: dataflow.thread.launch\n// CHECK-NOT: dataflow.thread private\n\nfunc.func @offset_strided(%buffer: memref) {\n %lower = arith.constant 5 : index\n %upper = arith.constant 9 : index\n %step = arith.constant 2 : index\n %value = arith.constant 7 : i32\n scf.forall (%index) = (%lower) to (%upper) step (%step) {\n memref.store %value, %buffer[%index] : memref\n }\n return\n}","why":"Non-normative evidence that a raw func-level scf.forall makes the selected stage fail with no output, so generated inputs place the effect-form forall inside the thread definition instead."},{"file_sha256":"7b2bea863a561d79b9cc6deb6f69e2c6e535d46e50f23d7fddb0e13d1a952c7d","kind":"example","lines":"1-57","path":"test/dataflow/unit/thread/valid.mlir","roles":["input_construction"],"text":"// RUN: loom %s | loom | FileCheck %s\n\n// Empty thread body carrying just the thread_ctrl slot per spec\n// section 5.4.1's `(args_*, thread_ctrl, iv_*)` layout.\n// CHECK-LABEL: dataflow.thread private @t_empty domain(#dataflow.thread_domain)() ctrl (%{{.*}}: none)\ndataflow.thread private @t_empty domain(#dataflow.thread_domain)() ctrl (%c: none) {\n dataflow.thread.yield\n}\n\n// Thread definition with two body operands, the thread_ctrl slot,\n// and one trailing grid iv slot.\n// CHECK-LABEL: dataflow.thread private @t_two_args domain(#dataflow.thread_domain)(%{{.*}}: i32, %{{.*}}: f32) ctrl (%{{.*}}: none) iv (%{{.*}}: index)\ndataflow.thread private @t_two_args domain(#dataflow.thread_domain)(%a: i32, %b: f32) ctrl (%c: none) iv (%i: index) {\n dataflow.thread.yield\n}\n\n// Every launch produces one completion token, including a launch carrying\n// mapped operands and a grid upper bound.\n// CHECK-LABEL: func.func @launch_demo\nfunc.func @launch_demo(%a: i32, %b: f32, %n: index) {\n // CHECK: %{{.*}} = dataflow.thread.launch @t_two_args(%{{.*}}, %{{.*}}) grid(%{{.*}}) : (i32, f32) -> !dataflow.thread_token\n %completion = dataflow.thread.launch @t_two_args(%a, %b) grid(%n) : (i32, f32) -> !dataflow.thread_token\n return\n}\n\n// Launch dependencies and waits both express unordered all-of completion.\n// CHECK-LABEL: func.func @wait_for_launches\nfunc.func @wait_for_launches() {\n // CHECK: %{{.*}} = dataflow.thread.launch @t_empty() : () -> !dataflow.thread_token\n %first = dataflow.thread.launch @t_empty() : () -> !dataflow.thread_token\n // CHECK: %{{.*}} = dataflow.thread.launch @t_empty() wait(%{{.*}}) : () -> !dataflow.thread_token\n %second = dataflow.thread.launch @t_empty() wait(%first) : () -> !dataflow.thread_token\n // CHECK: dataflow.thread.wait %{{.*}}, %{{.*}} : !dataflow.thread_token, !dataflow.thread_token\n dataflow.thread.wait %first, %second : !dataflow.thread_token, !dataflow.thread_token\n return\n}\n\n// Completion frontiers carry zero or more unordered none-typed values.\n// CHECK-LABEL: dataflow.thread private @t_frontier domain(#dataflow.thread_domain)() ctrl (%{{.*}}: none)\ndataflow.thread private @t_frontier domain(#dataflow.thread_domain)() ctrl (%ctrl: none) {\n // CHECK: dataflow.thread.yield %{{.*}} : none\n dataflow.thread.yield %ctrl : none\n}\n\n// A dynamic-work definition designates one ordinary input as its root payload\n// and has no coordinate suffix or launch extents.\n// CHECK-LABEL: dataflow.thread private @t_dynamic domain(#dataflow.thread_domain)\ndataflow.thread private @t_dynamic domain(#dataflow.thread_domain)(%work: i32) ctrl (%ctrl: none) {\n dataflow.thread.yield\n}\n\n// CHECK-LABEL: func.func @launch_dynamic\nfunc.func @launch_dynamic(%root: i32) {\n // CHECK: dataflow.thread.launch @t_dynamic(%{{.*}}) : (i32) -> !dataflow.thread_token\n %completion = dataflow.thread.launch @t_dynamic(%root) : (i32) -> !dataflow.thread_token\n return\n}","why":"Accepted spellings of dataflow.thread definitions (dense and dynamic domains, ctrl/iv clauses, thread.yield) and of dataflow.thread.launch with grid and wait clauses; used only for concrete syntax."},{"file_sha256":"3e2863717e0d5713893e5cabfb8e646c854432da0abd1478ee022ec7a610e62f","kind":"example","lines":"80-93","path":"test/raise/scf-to-dfg-atomic-publication.mlir","roles":["input_construction"],"text":"dataflow.thread private @parallel_channel_sender domain(#dataflow.thread_domain)(\n %channel: !dataflow.channel, %message: i32) ctrl (%start: none) {\n \"loom.spatial_region\"(%message, %channel)\n <{operandSegmentSizes = array,\n resultSegmentSizes = array}> ({\n ^bb0(%payload: i32, %output: !dataflow.channel):\n scf.forall (%lane) in (2) {\n dataflow.channel.send %output, %payload : !dataflow.channel\n }\n \"loom.spatial_yield\"()\n <{operandSegmentSizes = array}> : () -> ()\n }) {graph_name = \"parallel_channel_graph\", source_maps = []} :\n (i32, !dataflow.channel) -> ()\n dataflow.thread.yield","why":"Accepted spelling of an effect-form `scf.forall (%lane) in (2)` nested inside a dataflow.thread carrier, matching the sampled input shape."}],"primary_bundle_sha256":"bf761224d3ed8af6b66873c51c5c09d0291c0db4ea2bfc275ad80b806290ed1e","project":"PolyArch/loom","revision":"48615bc5925ef4b9db8b4550b5d4322933cf4b7b","schema":"spectriad.authoring-context/v1","selection_sha256":"b79feafca278e758c042da43242e7412169f3c7faee3662d7b43252e46b7c844"}