Provisional Research Note
Toward a Decision-Relative Definition of Evidence Strength
AIM Research Institute · APRILE Inc.
Status: PROVISIONAL / UNVERIFIED
Recorded: 11 August 2026
1. Purpose
This Research Note records a provisional attempt to further examine one of the open questions identified in the preceding research note on Evidence–Execution Proportionality:
How can Evidence Strength be measured?
The preceding note provisionally represented the relationship between evidence, hypothesis confidence, context, and autonomous execution as:
(Imax, Pmax) = F(E, H, C)
where:
-
E = Evidence Strength
-
H = Hypothesis Confidence
-
C = Context
-
Imax = maximum permissible Execution Intensity
-
Pmax = maximum permissible Execution Impact
The present note does not modify or validate that candidate relationship.
Its purpose is narrower:
What does it mean for evidence to be “strong” in relation to a specific hypothesis and decision?
This note records a provisional structure that emerged through repeated hypothetical case analysis.
It is not an established mathematical model, an established component of AIM, or a validated measurement framework.
2. The Initial Assumption
A natural starting point is to treat Evidence Strength primarily as a function of source credibility.
For example:
E = f(Source Reliability, Information Quantity, Recency)
This approach is intuitive.
However, repeated case analysis suggested that source credibility alone may be insufficient.
Information may be highly credible and still provide weak support for a particular hypothesis.
Conversely, some uncertainty may remain while the available evidence is already sufficient to support a particular decision.
This suggests a distinction between:
How credible is this information?
and
How strongly does this evidence support the specific hypothesis or decision being evaluated?
These may not be equivalent questions.
3. Evidence Strength May Be Hypothesis-Relative
The same evidence may have different significance depending on the hypothesis being evaluated.
For example, observing that three customers reduced orders may appear significant.
However, its evidential significance changes materially depending on whether the company has:
-
three customers,
-
ten customers,
-
or five hundred customers.
The absolute observation remains the same.
Its structural significance does not.
This suggests:
Evidence Strength ≠ Fixed Property of Evidence
and provisionally:
E = E(ℰ | H, C, t)
where:
-
ℰ = Evidence set
-
H = Hypothesis
-
C = Decision Context
-
t = evaluation time
-
Evidence Strength may therefore need to be evaluated relative to the hypothesis, context, and time under examination.
4. Source Credibility May Be Necessary but Insufficient
During case analysis, a distinction emerged between the identity of a source and the evidential value of the information it provides.
An official or identifiable source may have clear accountability while still possessing incentives that affect how information is presented.
Similarly, a named individual may provide information that is directly observable but incomplete.
This suggests that fixed credibility scores based solely on source categories may be inadequate.
For example:
Official Source ≠ Neutral Source
and:
Identifiable Source ≠ Sufficient Evidence
A provisional evaluation may therefore need to consider not only source identity, but also:
-
accountability,
-
directness,
-
independent corroboration,
-
information-generation pathway,
-
and source incentives.
This candidate dimension is provisionally referred to here as:
Verification Confidence (V)
5. Structural Relevance
A second dimension emerged repeatedly.
Evidence may be verified and factually correct while being too small relative to the evaluated structure to materially affect the hypothesis.
For example:
Three affected customers may be critical in a customer base of three, but close to immaterial in a customer base of five hundred.
Likewise, an observed change in a business unit representing 0.5% of total operations may not provide meaningful evidence about the condition of the entire organization.
This suggests a candidate dimension:
Structural Relevance (R)
Structural Relevance concerns whether the magnitude, proportion, functional importance, or structural position of an observation is sufficient to matter to the hypothesis being evaluated.
Thus:
Verified Evidence ≠ Structurally Relevant Evidence
6. Temporal Applicability
Evidence also appears to require a temporal relationship to the hypothesis.
A figure may be accurate but refer to a period before a significant change occurred.
Likewise, a current observation does not automatically justify a longer-term projection.
This suggests a distinction between factual accuracy and temporal applicability.
A candidate dimension is therefore:
Temporal Applicability (T)
The relevant questions include:
-
When was the evidence observed?
-
What period does the hypothesis concern?
-
Have subsequent actions already affected the reported figure?
-
Is there a reasonable basis for assuming continuity across the relevant time horizon?
-
This implies:
Historical Accuracy ≠ Current Applicability
and:
Current Observation ≠ Future Persistence
7. Structural Traceability
A fourth dimension emerged from cases in which an observed fact could plausibly arise from multiple underlying structures.
For example, a bank stopping additional lending may be consistent with deteriorating financial performance.
But it may also result from covenant violations, misuse of funds, collateral issues, concentration limits, or other causes.
The observation itself may be verified.
Its causal interpretation may remain unresolved.
This suggests a candidate dimension:
Structural Traceability (Tr)
Structural Traceability refers provisionally to:
The extent to which the material structural pathway connecting an observed piece of evidence to the hypothesis can be followed sufficiently for the decision under consideration.
This does not necessarily require identification of every root cause.
The relevant stopping point may instead be reached when remaining unknowns can no longer materially change the decision currently being evaluated.
Thus:
Decision Sufficiency ≠ Complete Causal Knowledge
8. A Provisional Evidence State
The purpose of separating these candidate dimensions is not to eliminate uncertainty or to imply that evidence can be made complete.
Rather, the purpose is to examine whether the available evidence is sufficiently verifiable, structurally relevant, temporally applicable, and traceable to support the hypothesis under evaluation without obscuring uncertainty that may still matter to the decision.
The preceding observations suggest that Evidence Strength may not initially be best represented as a single scalar score.
A provisional representation is:
E⃗ = (V, R, T, Tr)
where:
-
V = Verification Confidence
-
R = Structural Relevance
-
T = Temporal Applicability
-
Tr = Structural Traceability
This is not proposed as a validated measurement vector.
No numerical scales, normalization rules, weights, thresholds, or aggregation functions have been established.
The purpose of the representation is only to preserve four dimensions that repeatedly appeared during case analysis.
A future scalar Evidence Strength function might take a form such as:
E = F_E(V, R, T, Tr)
but the functional form of F_E remains undefined.
9. Evidence Strength and Decision Sufficiency May Be Different
A further distinction emerged during the analysis.
Incomplete evidence does not necessarily imply that a decision cannot yet be made.
Consider a financial case in which an uncertain additional cash inflow may or may not occur.
If the organization remains in the same material liquidity-risk state even under the most favorable plausible outcome, resolving that uncertainty may not change the relevant decision.
This suggests:
Evidence Completeness ≠ Decision Sufficiency
The relevant question may instead be:
Can any remaining plausible uncertainty still change the decision that matters?
This shifts the focus from eliminating uncertainty to identifying decision-relevant uncertainty.
10. Decision-Critical Variables
Not every unknown variable appears to require resolution.
A variable becomes relevant when its plausible variation could materially alter:
-
the hypothesis,
-
the decision,
-
the action,
-
the scale of the action,
-
or its timing.
-
Such variables are provisionally referred to here as:
Decision-Critical Variables (DCVs)
Let unresolved Decision-Critical Variables be:
U = {u₁, u₂, …, uₙ}
For each unresolved variable uᵢ, let Ωᵢ represent the set or range of states that remain reasonably plausible given the available evidence.
The combined plausible uncertainty space may then be represented as:
Ωᵤ = Ω₁ × Ω₂ × … × Ωₙ
This representation remains provisional.
11. Decision Robustness
A potentially important pattern appeared repeatedly in the case analysis.
If every reasonably plausible state of the unresolved variables leads to the same material decision, then the decision may already be sufficiently robust even though uncertainty remains.
Provisionally:
DR = 1 if and only if |{D(C, E⃗, u) : u ∈ Ωᵤ}| = 1
In plain language:
A decision is robust when all reasonably plausible unresolved states lead to the same material Decision State.
This is not a claim that the underlying evidence is complete.
It means only that the remaining uncertainty does not currently alter the relevant decision.
12. Decision Robustness and Action Robustness May Differ
Another distinction emerged.
The decision may remain unchanged while the specific action should still vary depending on unresolved information.
For example, a manufacturer may already know that delivery delays are unavoidable.
That decision state may be robust.
However, the number of delayed units, communication to customers, or revised delivery dates may depend on whether an external production option succeeds.
Thus:
Decision Robustness ≠ Action Robustness
If A represents an action function, Action Robustness may provisionally be represented as:
AR = 1 if and only if |{A(C, E⃗, u) : u ∈ Ωᵤ}| = 1
This distinction may allow a system to recognize:
“The decision is sufficiently supported, but the specific action is not yet sufficiently determined.”
13. Structural Recovery and Temporal Deferral
A further distinction became visible in financial cases.
An intervention may improve an immediate observable state without changing the structure generating the problem.
Examples may include:
-
additional borrowing,
-
temporary payment deferral,
-
accelerated receivable collection,
-
asset sales,
-
or other temporary liquidity measures.
These interventions may extend the time before failure without altering the underlying structure.
This suggests separating:
Structural Recovery
from:
Temporal Deferral
Conceptually:
Temporary Improvement ≠ Structural Recovery
A favorable short-term result should therefore not automatically be interpreted as evidence that the underlying problem has been corrected.
14. Probability May Not Always Be the First Question
The analysis also suggested that probability estimation may not always be necessary at the beginning of the evaluation.
If all reasonably plausible scenarios remain on the same side of the relevant Decision Boundary, estimating the precise probability of each scenario may not materially change the decision.
Conceptually:
No Decision Boundary Crossing → Detailed Probability Estimation May Be Unnecessary
If plausible scenarios cross the Decision Boundary, probability estimation, additional evidence acquisition, or scenario refinement may become relevant.
This suggests a possible sequence:
Plausible Scenarios → Boundary Test → Probability if Required
rather than automatically beginning with probability estimation.
15. Uncertainty Does Not Automatically Imply Conservatism
A significant issue emerged when considering unresolved uncertainty that cannot be eliminated before a decision deadline.
One possible approach would be to instruct the system to automatically select the “safer” option.
The present analysis does not support adopting that rule by default.
What constitutes the safer option may itself depend on:
-
objective,
-
priority,
-
role,
-
authority,
-
affected parties,
-
time horizon,
-
and competing forms of impact.
Therefore:
Unresolved Uncertainty ≠ Automatic Conservative Decision
unless such a decision rule has been explicitly established by the authorized human decision-maker.
Where a Decision-Critical Gap remains unresolved and cannot be resolved before the required decision point, the system may instead need to externalize:
-
what remains unknown,
-
why it matters,
-
which Decision Boundary it could cross,
-
what plausible states remain,
-
and why the evidence is not sufficient to resolve the decision autonomously.
The unresolved decision can then be returned to Human Judgment.
16. Human Judgment Does Not Retroactively Create Evidence Sufficiency
If a human chooses to proceed despite unresolved Decision-Critical uncertainty, that decision should not alter the recorded evidential state.
Thus:
Human Decision to Proceed ≠ Evidence Sufficiency
A record may therefore legitimately contain:
Decision Sufficiency = Not Ready
while also recording:
Human Decision = Proceed
This distinction may be important for later evaluation.
If the action eventually succeeds, the success should not retrospectively convert the original evidence into sufficient evidence.
This remains consistent with the preceding proposition:
Execution Success ≠ Retrospective Validation
17. Resolvable and Irreducible Decision Gaps
Unresolved Decision-Critical Gaps may also need to be separated according to whether they can reasonably be resolved before the required decision point.
Resolvable Gap
A Decision-Critical Gap for which relevant evidence can reasonably be obtained before the Decision Deadline.
Irreducible Gap
A Decision-Critical Gap that cannot reasonably be resolved before the Decision Deadline, including cases in which the relevant information does not yet exist or cannot yet be observed.
This distinction may support different system behavior:
Resolvable Gap → Evidence Acquisition
Irreducible Gap → Explicit Uncertainty Record → Human Judgment
This does not imply that Human Judgment resolves the evidential uncertainty.
It means that the limit of autonomous evidential evaluation has been reached under the current conditions.
18. A Provisional Architecture
The emerging candidate architecture can be represented as:
Decision Context → Decision Boundary → Hypothesis → Evidence Evaluation → Decision-Critical Variables → Plausible Uncertainty Space → Robustness Evaluation → Sufficiency → Human Judgment → Execution
Execution results may then return to the evidential pathway:
Executionₜ → Resultₜ → Evidenceₜ₊₁ → Hypothesis Reassessment
This architecture remains provisional.
19. Relation to the Previous Research Note
The preceding Research Note provisionally represented:
(Imax, Pmax) = F(E, H, C)
The present note examines only one part of that candidate relationship:
E — Evidence Strength
Specifically:
What structure may be required to evaluate Evidence Strength before it is used to support a hypothesis and subsequently constrain autonomous execution?
The broader research direction is therefore not to define Evidence Strength in isolation.
If a sufficiently specified account of E can eventually be developed, a subsequent question is whether and how Evidence
Strength relates to Hypothesis Confidence:
E → H ?
Only after that relationship is examined does the preceding candidate relationship return as a further research question:
(E, H, C) → (Imax, Pmax) ?
Accordingly, the provisional research pathway can currently be represented as:
Evidence Strength (E) → Hypothesis Confidence (H) → Permissible Execution Intensity / Impact
under Context (C)
This pathway is a research direction, not an established causal or mathematical relationship.
The present note does not define:
E → H
nor does it define:
(E, H, C) → (Imax, Pmax)
Both relationships remain open research questions.
20. Candidate Propositions
The present analysis produces the following provisional propositions for further testing:
-
Evidence Strength may be hypothesis-, context-, and time-relative rather than an intrinsic fixed property of information.
-
Source credibility alone may be insufficient to determine Evidence Strength.
-
Verified evidence may still lack Structural Relevance.
-
Accurate evidence may lack Temporal Applicability to the hypothesis under evaluation.
-
Structural Traceability may be required to determine what an observation actually supports.
-
Evidence completeness and Decision Sufficiency are not equivalent.
-
Not every unknown is Decision-Critical.
-
A decision may be robust even while material uncertainty remains.
-
Decision Robustness and Action Robustness may differ.
-
Temporary improvement should not automatically be interpreted as Structural Recovery.
-
Probability estimation may be unnecessary when no plausible unresolved scenario crosses the relevant Decision Boundary.
-
Unresolved uncertainty should not automatically be converted into a conservative decision unless such a rule has been explicitly authorized.
-
Human Judgment does not retroactively create Evidence Sufficiency.
-
Execution success does not retrospectively validate the evidence available when the original decision was made.
All propositions remain provisional and require further verification.
21. Open Research Questions
The present analysis leaves several questions unresolved.
How should Verification Confidence be measured?
How should Structural Relevance be normalized across domains?
How should Temporal Applicability be represented?
How should Structural Traceability be operationalized without requiring infinite causal investigation?
How should Decision-Critical Variables be identified consistently?
How should the plausible uncertainty space Ωᵤ be bounded?
What constitutes a material change in a Hypothesis, Decision, or Action?
Can Decision Robustness and Action Robustness be measured continuously rather than as binary states?
How should Evidence Strength relate to Hypothesis Confidence?
Under what conditions should Probability estimation be activated?
How should Decision Impact and irreversibility affect Human Judgment requirements?
Which parts of this evaluation should be performed by the acting AI, an independent monitoring layer, or humans?
How does this candidate architecture compare with existing work in decision theory, robust decision-making, uncertainty quantification, AI Safety, AI Control, and human oversight?
These questions remain open.
22. Current Research Position
The present analysis does not establish a validated formula for Evidence Strength.
It does not establish that:
E⃗ = (V, R, T, Tr)
is complete.
It does not establish numerical scales, weighting rules, thresholds, aggregation functions, or domain-independent normalization methods.
It also does not establish Decision Robustness or Action Robustness as formal AIM components.
The narrower result of the present analysis is:
Evidence Strength may require evaluation not only of whether information is credible, but whether it is structurally relevant, temporally applicable, and sufficiently traceable to the hypothesis under evaluation.
A second provisional result is:
The sufficiency of evidence for a decision may depend less on whether all uncertainty has been eliminated than on whether any remaining plausible uncertainty can still materially change the decision.
These propositions require further testing.
23. Status
PROVISIONAL / UNVERIFIED
The variables, representations, definitions, and relationships recorded in this note are candidate research constructs.
They are not established components of AIM.
No claim of novelty is made.
Future work may include:
-
additional cross-domain case testing,
-
formal definition of Evidence Strength,
-
examination of the relationship between Evidence Strength and Hypothesis Confidence,
-
mathematical treatment of Decision and Action Robustness,
-
comparison with existing decision theory and robust decision-making literature,
-
retrospective testing against real-world decision records,
-
simulation with autonomous AI agents,
-
and eventual examination of how Evidence Strength and Hypothesis Confidence should constrain Execution Intensity and Execution Impact.
If subsequent testing contradicts any of the candidate constructs recorded here, those contradictions should be preserved as part of the research record.
暫定研究ノート
意思決定に相対的な「エビデンス強度」の定義に向けて
エビデンス評価・意思決定上重要な不確実性・判断の頑健性に関する暫定研究ノート
AIM Research Institute · APRILE Inc.
ステータス:暫定/未検証(PROVISIONAL / UNVERIFIED)
記録日:2026年8月11日
1. 目的
本研究ノートは、先行する「エビデンスと実行の比例関係」に関する研究ノートで残された未解決の問いの一つを、さらに検討するための暫定的な記録である。
その問いは、
エビデンスの強さは、どのように評価できるのか?
である。
先行研究ノートでは、エビデンス、仮説への確信度、文脈、自律的な実行の関係を、暫定的に次のように表した。
(Imax, Pmax) = F(E, H, C)
ここで、
-
E = エビデンス強度
-
H = 仮説への確信度
-
C = 文脈
-
Imax = 許容される実行強度の最大値
-
Pmax = 許容される実行影響度の最大値
を意味する。
本ノートは、この候補関係を変更したり、正しいと証明したりするものではない。
今回扱う問いは、より限定的である。
特定の仮説や意思決定との関係において、エビデンスが「強い」とは何を意味するのか?
本ノートでは、複数の仮想ケースを繰り返し分析する過程で現れた暫定的な構造を記録する。
これは確立された数理モデルではなく、AIMの正式な構成要素でもなく、検証済みの測定手法でもない。
2. 当初の仮定
自然な出発点の一つは、エビデンス強度を主として「情報源の信頼性」によって評価することである。
たとえば、
E = f(情報源の信頼性、情報量、鮮度)
のような考え方である。
これは直感的には理解しやすい。
しかし、ケース分析を繰り返す中で、情報源の信頼性だけでは不十分である可能性が見えてきた。
情報そのものの信頼性が高くても、特定の仮説を支持するエビデンスとしては弱い場合がある。
反対に、一定の不確実性が残っていても、特定の意思決定を行うには、すでに十分なエビデンスが存在している場合もある。
したがって、次の二つは区別する必要がある可能性がある。
この情報はどの程度信用できるのか?
と、
このエビデンスは、評価対象となっている特定の仮説や意思決定を、どの程度強く支持しているのか?
である。
この二つは同じ問いではない可能性がある。
3. エビデンス強度は仮説に相対的である可能性
同じエビデンスでも、評価対象となる仮説によって、その意味は異なり得る。
たとえば、「3社の顧客が発注量を減らした」という観測があったとする。
一見すると重要な情報に見える。
しかし、その企業の顧客数が、
-
3社なのか
-
10社なのか
-
500社なのか
によって、その情報が持つ意味は大きく変わる。
観測された事実そのものは変わらない。
しかし、その事実が全体構造の中で持つ意味は変わる。
したがって、
エビデンス強度 ≠ エビデンスそのものに固定された属性
である可能性がある。
暫定的には、
E = E(ℰ | H, C, t)
と表すことができる。
ここで、
-
ℰ = エビデンス集合
-
H = 評価対象となる仮説
-
C = 意思決定の文脈
-
t = 評価時点
である。
つまり、エビデンス強度は情報単体で決まるのではなく、何について、どのような状況で、いつ判断しようとしているのかとの関係で評価する必要がある可能性がある。
4. 情報源の信頼性は、必要であっても十分ではない可能性
ケース分析では、情報源の身元と、その情報が持つエビデンスとしての価値を分けて考える必要性が現れた。
公的機関や身元が明確な情報源には、責任主体を追跡できるという特徴がある。
しかし、その情報の提示方法に影響する利害や動機が存在する場合もある。
同様に、実名の人物による情報であっても、その人物が観測できる範囲が限定されていれば、情報は不完全である可能性がある。
したがって、
公的情報源 ≠ 中立的情報源
であり、
身元が明確な情報源 ≠ 十分なエビデンス
でもある。
暫定的な評価では、情報源の種類だけではなく、
-
誰がその情報に責任を持っているのか
-
一次情報なのか、伝聞なのか
-
独立した別経路から確認できるのか
-
その情報はどのような過程で生成されたのか
-
情報提供者にどのような利害があるのか
などを見る必要がある可能性がある。
本ノートでは、この候補次元を暫定的に、
検証確度(Verification Confidence:V)
と呼ぶ。
5. 構造的関連性
第二の次元として繰り返し現れたのが、構造的関連性(Structural Relevance:R)である。
エビデンスが事実として正確であっても、評価対象となる構造全体に対して規模が小さすぎれば、その仮説を評価するうえで実質的な意味を持たない可能性がある。
たとえば、3社の顧客への影響は、顧客総数が3社なら極めて重大である。
しかし、顧客総数が500社なら、企業全体の状態を判断するエビデンスとしては弱い可能性がある。
同様に、企業全体の0.5%しか占めない部門で観測された変化が、そのまま企業全体の状態を示すとは限らない。
構造的関連性とは、観測された事象の、
-
規模
-
全体に占める割合
-
機能的重要性
-
構造上の位置
などが、評価対象となる仮説に対して十分な意味を持つかどうかを見るものである。
したがって、
検証されたエビデンス ≠ 構造的に重要なエビデンス
である。
6. 時間的適用可能性
第三の候補次元は、時間的適用可能性(Temporal Applicability:T)である。
ある数値が正確であっても、重要な変化が起こる前の期間を示している場合がある。
また、現在観測されている状態が、そのまま中長期的に継続するとは限らない。
したがって、「事実として正しいこと」と「現在の仮説に時間的に適用できること」は分ける必要がある可能性がある。
ここでは、たとえば以下を確認する。
-
そのエビデンスはいつ観測されたのか
-
仮説はどの期間を対象としているのか
-
観測後に状況を変える行動がすでに行われていないか
-
現在の状態が対象期間まで継続すると考える合理的根拠があるか
したがって、
過去の正確性 ≠ 現在への適用可能性
であり、
現在の観測 ≠ 将来も継続することの証明
である。
7. 構造的追跡可能性
第四の候補次元は、構造的追跡可能性(Structural Traceability:Tr)である。
これは、一つの観測事実が複数の異なる原因や構造から生じ得るケースから現れた。
たとえば、銀行が追加融資を停止したという事実が確認されたとする。
これは業績悪化によるものかもしれない。
しかし、
-
資金使途の問題
-
財務上の契約条件への抵触
-
担保上の問題
-
融資集中限度
-
その他の銀行側の判断
によって発生した可能性もある。
つまり、「銀行が融資を止めた」という事実が確認できても、
なぜ止めたのか
が分からなければ、その事実がどの仮説を支持するのかは確定しない。
本ノートでは、構造的追跡可能性を暫定的に、
観測されたエビデンスから評価対象となる仮説までを結ぶ、意思決定上重要な構造的経路を、その判断に必要な範囲まで追跡できる程度
と考える。
ただし、すべての根本原因を完全に特定する必要があるとは限らない。
残っている不明点が合理的な範囲で変化しても、現在評価している意思決定を実質的に変更しなくなった地点で、追跡を停止できる可能性がある。
したがって、
意思決定に十分な状態 ≠ 原因を完全に把握した状態
である。
8. 暫定的なエビデンス状態
これら4つの候補次元を分離する目的は、不確実性を完全に除去することではない。
また、エビデンスを完全なものにできると仮定することでもない。
目的は、利用可能なエビデンスが、
-
十分に検証可能か
-
構造的に意味を持つか
-
対象期間に適用できるか
-
仮説までの重要な経路を追跡できるか
を確認すると同時に、意思決定を変え得る不確実性を覆い隠さないことである。
ここまでの分析から、エビデンス強度を最初から一つの点数にまとめることは適切ではない可能性がある。
暫定的には、
E⃗ = (V, R, T, Tr)
と表現する。
ここで、
-
V = 検証確度
-
R = 構造的関連性
-
T = 時間的適用可能性
-
Tr = 構造的追跡可能性
である。
これは検証済みの測定ベクトルではない。
現時点では、
-
数値尺度
-
正規化方法
-
重み
-
閾値
-
統合方法
のいずれも確立していない。
この表現の目的は、ケース分析で繰り返し現れた4つの次元を保存することにある。
将来的には、
E = F_E(V, R, T, Tr)
のように、一つのエビデンス強度へ統合できる可能性もある。
ただし、F_Eの具体的な形は未定義である。
9. エビデンス強度と「意思決定に十分か」は同じではない
分析の過程で、さらに重要な区別が現れた。
エビデンスが完全ではないことは、必ずしも「まだ意思決定できない」ことを意味しない。
たとえば、企業に追加の資金流入があるかどうかが分からないとする。
しかし、合理的に考えられる最も良い結果が起きても、企業が同じ重大な資金繰りリスクの状態に留まるのであれば、その不確実性を解消しても判断は変わらない。
したがって、
エビデンスの完全性 ≠ 意思決定に十分であること
である可能性がある。
ここで重要となる問いは、
残っている合理的な不確実性の中に、現在重要な意思決定をなお変更し得るものが存在するか?
である。
つまり、すべての不確実性をなくすことではなく、判断を変え得る不確実性を特定することが重要になる。
10. 意思決定上重要な変数
すべての不明点を解消する必要があるとは限らない。
ある変数が合理的な範囲で変化したときに、
-
仮説
-
意思決定
-
具体的な行動
-
行動の規模
-
行動の時期
のいずれかを実質的に変更し得る場合、その変数は重要である。
本ノートでは、このような変数を暫定的に、
意思決定上重要な変数(Decision-Critical Variables:DCVs)
と呼ぶ。
未解消の意思決定上重要な変数を、
U = {u₁, u₂, …, uₙ}
とする。
各変数 uᵢ について、現在のエビデンスから合理的に成立し得る状態の範囲を Ωᵢ とする。
すると、未解消の変数全体が取り得る合理的な状態の範囲は、
Ωᵤ = Ω₁ × Ω₂ × … × Ωₙ
と暫定的に表現できる。
この表現も、現時点では候補である。
11. 意思決定の頑健性
ケース分析では、重要なパターンが繰り返し現れた。
未解消の変数について、合理的に成立し得る状態をすべて考えても、到達する重要な意思決定が同じである場合がある。
この場合、不確実性が残っていても、その意思決定は十分に頑健(Robust)である可能性がある。
暫定的には、
DR = 1 ⇔ |{D(C, E⃗, u) : u ∈ Ωᵤ}| = 1
と表現する。
意味は単純である。
合理的に残っている不確実性を動かしても、到達する重要な意思決定が一つしかないなら、その意思決定は頑健である。
これは、エビデンスが完全であることを意味しない。
残っている不確実性が、現在の意思決定を変更しないという意味である。
12. 意思決定の頑健性と行動の頑健性は異なり得る
さらに、意思決定と具体的な行動を分離する必要性が現れた。
意思決定そのものは変わらなくても、未解消の情報によって具体的な行動を変える必要がある場合がある。
たとえば製造業において、納期遅延が避けられないこと自体はすでに明確になっているとする。
その判断は変わらない。
しかし、
-
何台遅れるのか
-
顧客へどのように説明するのか
-
新しい納期をいつにするのか
は、外注生産の結果によって変わる可能性がある。
したがって、
意思決定の頑健性 ≠ 行動の頑健性
である。
行動を A とすると、暫定的には、
AR = 1 ⇔ |{A(C, E⃗, u) : u ∈ Ωᵤ}| = 1
と表現できる。
この区別によって、
「判断そのものは十分に支持されているが、具体的な行動はまだ確定できない」
という状態を識別できる可能性がある。
13. 構造的回復と時間的先送り
財務ケースでは、別の重要な区別も現れた。
何らかの介入によって直近の状態が改善しても、問題を生み出している構造そのものが変化していない場合がある。
たとえば、
-
追加借入
-
支払期限の一時的な延期
-
売掛金回収の前倒し
-
資産売却
-
その他の一時的な資金対策
などである。
これらによって資金ショートまでの時間が延びても、問題を生み出している構造自体が変わっていなければ、回復とは限らない。
そこで、
構造的回復(Structural Recovery)
と、
時間的先送り(Temporal Deferral)
を分ける必要がある可能性がある。
つまり、
一時的な改善 ≠ 構造的な回復
である。
短期的に良い結果が出たことだけをもって、根本的な問題が解消されたと判断すべきではない。
14. 確率は必ずしも最初に計算する必要はない
今回の分析では、確率推定を必ず最初に行う必要はない可能性も見えてきた。
合理的に成立し得るすべてのシナリオが、意思決定の境界線の同じ側にある場合、それぞれのシナリオの正確な確率を計算しても、意思決定は変わらない可能性がある。
したがって、
意思決定境界を跨ぐシナリオがない → 詳細な確率推定は不要である可能性
がある。
反対に、合理的なシナリオの中に意思決定境界を跨ぐものが存在する場合には、
-
確率推定
-
追加エビデンスの取得
-
シナリオ範囲の精緻化
などが重要になる可能性がある。
したがって、候補となる評価順序は、
合理的シナリオ → 境界判定 → 必要な場合のみ確率推定
である。
15. 不確実性が残ることは、自動的に「安全側」を選ぶことを意味しない
意思決定期限までに解消できない不確実性について検討する中で、重要な問題が現れた。
不確実性が残る場合、AIが自動的に「より安全な選択肢」を選ぶという設計も考えられる。
しかし、今回の分析では、そのルールを初期設定として採用する根拠は得られていない。
なぜなら、「何が安全か」自体が、
-
目的
-
優先順位
-
役割
-
権限
-
影響を受ける主体
-
時間軸
-
複数の異なる影響
によって変わるからである。
したがって、
未解消の不確実性 ≠ 自動的に保守的な意思決定
である。
ただし、権限を持つ人間の意思決定者が、そのような判断ルールを事前に明示している場合は別である。
意思決定上重要な不明点が残り、それを期限までに解消できない場合、システムが行うべきことは、自動的に「安全側」を選ぶことではなく、
-
何が分かっていないのか
-
なぜそれが重要なのか
-
どの意思決定境界を跨ぐ可能性があるのか
-
どのような合理的状態が残っているのか
-
なぜ現在のエビデンスだけでは自律的に判断できないのか
を明示することである可能性がある。
そのうえで、人間の判断へ返す。
16. 人間の判断は、エビデンスの十分性を後から作り出さない
意思決定上重要な不確実性が残った状態で、人間が「実行する」と判断したとしても、その判断によって元のエビデンス状態が変更されるべきではない。
したがって、
人間が実行を決定したこと ≠ エビデンスが十分だったこと
である。
記録上、
意思決定に必要なエビデンス = 未充足
でありながら、
人間の判断 = 実行
という状態は成立し得る。
この区別は、その後の検証において重要になる可能性がある。
最終的に実行が成功したとしても、その成功によって、当初利用可能だったエビデンスが遡及的に「十分だった」ことにはならない。
これは先行研究ノートで提示した、
実行成功 ≠ 元のエビデンスの遡及的正当化
という命題とも整合する。
17. 解消可能な不明点と、期限内には解消できない不明点
未解消の意思決定上重要な不明点は、意思決定期限までに合理的に解消できるかどうかによって分ける必要がある可能性がある。
・解消可能な不明点
意思決定期限までに、必要なエビデンスを合理的に取得できるもの。
・期限内には解消できない不明点
意思決定期限までに合理的に解消できないもの。
必要な情報がまだ存在しない場合や、その時点では観測できない場合も含む。
この区別によって、
解消可能な不明点 → 追加エビデンスを取得
期限内には解消できない不明点 → 不確実性を明示して記録 → 人間の判断
という異なる経路を設定できる可能性がある。
これは、人間が判断することによって不確実性そのものが解消されるという意味ではない。
現在の条件下で、AIによるエビデンス評価の限界に到達したことを意味する。
18. 暫定的な全体構造
ここまでに現れた候補構造は、暫定的に次のように表現できる。
意思決定の文脈 → 意思決定境界 → 仮説 → エビデンス評価 → 意思決定上重要な変数 → 合理的な不確実性の範囲 → 頑健性評価 → 十分性評価 → 人間の判断 → 実行
そして、実行によって得られた結果は、再びエビデンス評価へ戻る。
実行ₜ → 結果ₜ → エビデンスₜ₊₁ → 仮説の再評価
この構造も現時点では暫定的である。
19. 先行研究ノートとの関係
先行研究ノートでは、
(Imax, Pmax) = F(E, H, C)
という候補関係を提示した。
今回検討しているのは、そのうちの一つである。
E = エビデンス強度
である。
具体的な問いは、
エビデンスが仮説を支持する材料として使用され、その先でAIの自律的な実行範囲を制約する前に、エビデンス強度を評価するにはどのような構造が必要なのか?
である。
したがって、研究の目的は、エビデンス強度だけを孤立して定義することではない。
将来的に E を十分に具体化できた場合、次の問いは、
エビデンス強度と仮説への確信度は、どのような関係にあるのか?
となる。
暫定的には、
E → H ?
である。
その関係を検討した後に、先行研究ノートの問い、
(E, H, C) → (Imax, Pmax) ?
へ戻る。
したがって、現在見えている研究経路は、
エビデンス強度(E) → 仮説への確信度(H) → 許容される実行強度・実行影響度
ただし、文脈(C)のもとで
と表現できる。
ただし、これは研究の方向を示したものであり、確立された因果関係や数理関係ではない。
本ノートでは、
E → H
をまだ定義していない。
また、
(E, H, C) → (Imax, Pmax)
も定義していない。
どちらも引き続き未解決の研究課題である。
20. 暫定命題
今回の分析から、今後さらに検証するための暫定命題として、以下を記録する。
-
エビデンス強度は、情報そのものに内在する固定属性ではなく、仮説、文脈、時間に相対的である可能性がある。
-
情報源の信頼性だけでは、エビデンス強度を決定するには不十分である可能性がある。
-
検証されたエビデンスであっても、構造的関連性を欠く可能性がある。
-
正確なエビデンスであっても、評価対象となる仮説に対する時間的適用可能性を欠く可能性がある。
-
観測された事実が実際に何を支持しているのかを判断するために、構造的追跡可能性が必要となる可能性がある。
-
エビデンスが完全であることと、意思決定に十分であることは同じではない。
-
すべての不明点が、意思決定上重要であるとは限らない。
-
重要な不確実性が残っていても、意思決定は頑健であり得る。
-
意思決定の頑健性と、具体的な行動の頑健性は異なり得る。
-
一時的な改善を、自動的に構造的回復と解釈すべきではない。
-
合理的に残るシナリオが意思決定境界を跨がない場合、詳細な確率推定は不要である可能性がある。
-
未解消の不確実性を、事前に明示された権限ルールなしに、自動的に「安全側」の意思決定へ変換すべきではない。
-
人間による意思決定は、元のエビデンスを遡及的に十分なものへ変えない。
-
実行が成功しても、意思決定時点で利用可能だったエビデンスが十分だったことの遡及的証明にはならない。
以上はいずれも暫定命題であり、今後の検証を必要とする。
21. 未解決の研究課題
今回の分析でも、複数の問いが未解決のまま残っている。
検証確度はどのように測定すべきか。
構造的関連性を、異なる領域間でどのように標準化できるか。
時間的適用可能性をどのように表現すべきか。
無限の原因追跡に陥ることなく、構造的追跡可能性をどのように実装できるか。
意思決定上重要な変数を、一貫した方法でどのように特定できるか。
合理的な不確実性の範囲 Ωᵤ をどこまでとするべきか。
仮説、意思決定、行動における「実質的な変化」をどのように定義するか。
意思決定の頑健性と行動の頑健性を、0か1ではなく連続的に評価できるか。
エビデンス強度と仮説への確信度は、どのような関係にあるのか。
どの条件で確率推定を開始すべきか。
意思決定の影響度や不可逆性は、人間判断を必要とする条件にどのように影響するのか。
この評価のどこまでを実行主体であるAIが行い、どこからを独立した監視層または人間が担うべきか。
この候補構造は、意思決定理論、頑健意思決定、不確実性定量化、AI Safety、AI Control、人間による監督に関する既存研究とどのような関係にあるのか。
これらはすべて、今後の研究課題として残る。
22. 現時点での研究上の位置づけ
今回の分析によって、エビデンス強度の検証済み数式が確立されたわけではない。
また、
E⃗ = (V, R, T, Tr)
が完全な構造であることも確認されていない。
数値尺度、重み付け、閾値、統合関数、領域を超えて使用できる標準化方法も確立していない。
意思決定の頑健性や行動の頑健性についても、AIMの正式な構成要素として確立したものではない。
現時点で記録できる、より限定的な結果は次の通りである。
エビデンス強度を評価するためには、情報が信用できるかだけではなく、その情報が構造的に関連しているか、対象期間に適用できるか、そして評価対象となる仮説までの重要な経路を十分に追跡できるかを検討する必要がある可能性がある。
さらに、第二の暫定的な結果として、
ある意思決定に対してエビデンスが十分かどうかは、すべての不確実性が解消されたかではなく、残っている合理的な不確実性の中に、その意思決定をなお実質的に変更し得るものが存在するかによって決まる可能性がある。
という構造が現れている。
これらはいずれも、今後さらに検証する必要がある。
23. ステータス
暫定/未検証(PROVISIONAL / UNVERIFIED)
本ノートに記録した変数、表現、定義、関係は、すべて研究上の候補構造である。
現時点では、AIMの確立された構成要素ではない。
新規性についての主張も行わない。
今後の研究として、以下が考えられる。
-
異なる領域を用いた追加ケース検証
-
エビデンス強度の正式な定義
-
エビデンス強度と仮説への確信度の関係の検討
-
意思決定の頑健性と行動の頑健性の数理的検討
-
既存の意思決定理論および頑健意思決定研究との比較
-
実際の意思決定記録を用いた遡及的検証
-
自律型AIエージェントを用いたシミュレーション
-
エビデンス強度と仮説への確信度が、許容される実行強度と実行影響度をどのように制約すべきかの検討
-
今後の検証によって、本ノートに記録した候補構造と矛盾する結果が得られた場合、その矛盾も研究記録の一部として保存する。
