{"id":{"repo_id":"embry-riddle","oai_identifier":"oai:commons.erau.edu:edt-2025"},"canonical_url":"https://search.dev.ndltd.org/etd/embry-riddle/oai:commons.erau.edu:edt-2025","repository":{"repo_id":"embry-riddle","name":"Embry Riddle Aeronautical University","base_url":"https://commons.erau.edu/do/oai/"},"display":{"title":"Study of Output and Behavior of LLMs Using Confidence Framing in Prompt Engineering","abstract":"<p>While prompt engineering is pivotal for shaping Large Language Model (LLM) outputs, the impact of confidence framing on behavioral calibration remains underexplored. This study investigates the ways in which psychological framing, utilizing techniques such as capability praise, role amplification, and doubt induction, affects linguistic tone, objective accuracy, and internal calibration. A 1,080-trial experimental matrix evaluated six diverse models across factual, logical, coding, and cyber security domains. Analysis using the Kruskal-Wallis H-test revealed highly significant behavioral shifts across all measured dimensions, providing conclusive evidence that the applied frames exert a substantial influence on model performance.</p> <p>The findings identify a distinct cognitive trade-off. While confidence-boosting language produced more assertive and fluent outputs, it significantly degraded factual reliability and internal calibration in larger proprietary models. However, a paradox was observed in small language models, where authoritative or more confident frames acted as a corrective focusing mechanism that improved calibration. In the cyber security domain, doubt-inducing frames successfully weaponized alignment guardrails, increasing aggregate refusal rates from 43.3% to 71.1%. These results suggest that linguistic confidence is a trailing indicator of internal alignment rather than a marker of latent truth. This work establishes that overconfident framing introduces critical vulnerabilities in factual and logical domains, while simultaneously offering a potential reliability boost for lightweight models in regulated enterprise environments.</p>","abstract_html":"&lt;p&gt;While prompt engineering is pivotal for shaping Large Language Model (LLM) outputs, the impact of confidence framing on behavioral calibration remains underexplored. This study investigates the ways in which psychological framing, utilizing techniques such as capability praise, role amplification, and doubt induction, affects linguistic tone, objective accuracy, and internal calibration. A 1,080-trial experimental matrix evaluated six diverse models across factual, logical, coding, and cyber security domains. Analysis using the Kruskal-Wallis H-test revealed highly significant behavioral shifts across all measured dimensions, providing conclusive evidence that the applied frames exert a substantial influence on model performance.&lt;/p&gt; &lt;p&gt;The findings identify a distinct cognitive trade-off. While confidence-boosting language produced more assertive and fluent outputs, it significantly degraded factual reliability and internal calibration in larger proprietary models. However, a paradox was observed in small language models, where authoritative or more confident frames acted as a corrective focusing mechanism that improved calibration. In the cyber security domain, doubt-inducing frames successfully weaponized alignment guardrails, increasing aggregate refusal rates from 43.3% to 71.1%. These results suggest that linguistic confidence is a trailing indicator of internal alignment rather than a marker of latent truth. This work establishes that overconfident framing introduces critical vulnerabilities in factual and logical domains, while simultaneously offering a potential reliability boost for lightweight models in regulated enterprise environments.&lt;/p&gt;","abstract_has_math":false,"creators":["Parrilla, Micah"],"institution":null,"degree_name":"Master of Science in Computer Science","degree_level":"Thesis - Open Access","degree_discipline":"Electrical Engineering and Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026-04-01T07:00:00Z","date_published":"2026-04-01T07:00:00Z","updated_at":"2026-07-27T19:26:22Z","subjects":["prompt engineering","large language models","confidence framing","calibration index","cyber security alignment","hallucinations","self-reported confidence","Artificial Intelligence and Robotics","Computer Engineering","Theory and Algorithms"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://commons.erau.edu/edt/979","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Parrilla, Micah"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical Engineering and Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis - Open Access"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science in Computer Science"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["prompt engineering","large language models","confidence framing","calibration index","cyber security alignment","hallucinations","self-reported confidence","Artificial Intelligence and Robotics","Computer Engineering","Theory and Algorithms"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://commons.erau.edu/edt/979"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>While prompt engineering is pivotal for shaping Large Language Model (LLM) outputs, the impact of confidence framing on behavioral calibration remains underexplored. This study investigates the ways in which psychological framing, utilizing techniques such as capability praise, role amplification, and doubt induction, affects linguistic tone, objective accuracy, and internal calibration. A 1,080-trial experimental matrix evaluated six diverse models across factual, logical, coding, and cyber security domains. Analysis using the Kruskal-Wallis H-test revealed highly significant behavioral shifts across all measured dimensions, providing conclusive evidence that the applied frames exert a substantial influence on model performance.</p> <p>The findings identify a distinct cognitive trade-off. While confidence-boosting language produced more assertive and fluent outputs, it significantly degraded factual reliability and internal calibration in larger proprietary models. However, a paradox was observed in small language models, where authoritative or more confident frames acted as a corrective focusing mechanism that improved calibration. In the cyber security domain, doubt-inducing frames successfully weaponized alignment guardrails, increasing aggregate refusal rates from 43.3% to 71.1%. These results suggest that linguistic confidence is a trailing indicator of internal alignment rather than a marker of latent truth. This work establishes that overconfident framing introduces critical vulnerabilities in factual and logical domains, while simultaneously offering a potential reliability boost for lightweight models in regulated enterprise environments.</p>"]},{"key":"dc:title","label":"Title","values":["Study of Output and Behavior of LLMs Using Confidence Framing in Prompt Engineering"]}]}],"canonical_facts":{"dc:creator":["Parrilla, Micah"],"dc:description.abstract":["<p>While prompt engineering is pivotal for shaping Large Language Model (LLM) outputs, the impact of confidence framing on behavioral calibration remains underexplored. This study investigates the ways in which psychological framing, utilizing techniques such as capability praise, role amplification, and doubt induction, affects linguistic tone, objective accuracy, and internal calibration. A 1,080-trial experimental matrix evaluated six diverse models across factual, logical, coding, and cyber security domains. Analysis using the Kruskal-Wallis H-test revealed highly significant behavioral shifts across all measured dimensions, providing conclusive evidence that the applied frames exert a substantial influence on model performance.</p> <p>The findings identify a distinct cognitive trade-off. While confidence-boosting language produced more assertive and fluent outputs, it significantly degraded factual reliability and internal calibration in larger proprietary models. However, a paradox was observed in small language models, where authoritative or more confident frames acted as a corrective focusing mechanism that improved calibration. In the cyber security domain, doubt-inducing frames successfully weaponized alignment guardrails, increasing aggregate refusal rates from 43.3% to 71.1%. These results suggest that linguistic confidence is a trailing indicator of internal alignment rather than a marker of latent truth. This work establishes that overconfident framing introduces critical vulnerabilities in factual and logical domains, while simultaneously offering a potential reliability boost for lightweight models in regulated enterprise environments.</p>"],"dc:identifier":["https://commons.erau.edu/edt/979"],"dc:subject":["prompt engineering","large language models","confidence framing","calibration index","cyber security alignment","hallucinations","self-reported confidence","Artificial Intelligence and Robotics","Computer Engineering","Theory and Algorithms"],"dc:title":["Study of Output and Behavior of LLMs Using Confidence Framing in Prompt Engineering"],"thesis:degree_discipline":["Electrical Engineering and Computer Science"],"thesis:degree_level":["Thesis - Open Access"],"thesis:degree_name":["Master of Science in Computer Science"]},"updated_at":"2026-07-27T19:26:22Z"}