Prompt Engineering Best Practices — AnyCompany Support Workshop

🎯 The 4 Pillars of Effective Prompts

80% of prompt quality comes from 4 fundamentals. Master these and every prompt you write will be dramatically better.

🎯

1. Clarity

Say exactly what you mean. If a colleague would ask "what do you mean?" — your prompt needs work.

📚

2. Context

Give the AI the background it needs: domain, data, situation, constraints. Without context, it guesses.

👤

3. Role Assignment

Tell the AI who to be. A "senior IRT TL briefing leadership" focuses on different signals than a "Live-chat agent drafting a Pax response."

📋

4. Output Framing

Define what "done" looks like — format, length, structure, style. No framing = unpredictable output.

📚 4 Types of Context

Context is the most impactful pillar for GS workflows. Skip any type and the output suffers in a specific way:

Type	What it tells the AI	If you skip it...	GS example
Domain	Industry, market, business area	Generic, non-specific answers	"In the context of GS Support for Southeast Asian ride-hailing..."
Data	Specific case data, transcripts, history	AI hallucinates plausible booking IDs / SOP steps	"Here is the D365 case + last 5 chat turns: [data]"
Situational	Why you need this now — the trigger	Wrong tone, urgency, severity framing	"True Safety case escalation" vs "Routine MIWI write-up"
Constraints	Rules, limits, requirements	Ignores your standards	"Default SGD; redact PAX/DAX names; cite SOP article ID"

👀 See the Difference

❌ Without 4 Pillars

Prompt: "Summarise this case"

The Pax seems to have an issue with their ride. They mentioned some concerns about the driver. The case looks like it might need follow-up. Overall this appears to be a customer service matter to resolve.

⚠️ Vague, no severity, no booking, no SOP citation, not stakeholder-ready, would fail QA audit

✅ With 4 Pillars

Prompt: Role + Context + Structure + Constraints

RISK RATING: 🟡 AMBER
Transaction Analysis: Volume surged 271% (4,200→15,600). Avg txn flat at $10 suggests organic growth.
Chargeback Alert: Rate 0.3%→4.1%, exceeding 1.0% benchmark by 4x.
Action: 1) Issue warning letter (Ops, 5 days) 2) Reduce PayLater limit (Risk, immediate)

✓ Structured, data-cited, specific actions with owners and timelines

💡

Same data, dramatically different quality. The difference is entirely in the prompt — not the AI model.

🧠 Chain-of-Thought Reasoning

Financial decisions require multi-step logic. CoT makes reasoning visible and auditable — the AI shows its work.

💬

Zero-Shot CoT

Just add "Think step by step." No examples needed. Best for quick calculations and simple logic.

📝

Few-Shot CoT

Provide one example with reasoning. The AI follows the exact same pattern. Best for consistent processes.

🔭

Step-Back

"First identify key factors, then analyze." Forces prioritization before writing. Best for complex analysis.

🗳️

Self-Consistency

"Solve 3 ways, report majority." Multiple approaches catch what a single analysis misses. Best for high-stakes.

👀 Zero-Shot CoT: Before & After

Question: "Is this IRT case true Safety P1 or downgradable?"

❌ Without CoT

Yes, this looks like a true Safety case. The Pax raised a concern that warrants IRT attention.

⚠️ No reasoning. No SOP citation. No severity criteria visible. Not QA-defensible. Could be wrong (= 30-min SLA loss + DSAT).

✅ With "Think step by step"

Pax message: "the driver smelled like beer when I got in"
→ Explicit impairment language? YES [SOP §8.3 trigger]
→ Ride concluded <1 hour ago? YES (22:14 vs 22:38)
→ Two or more P1 indicators? YES → P1 confirmed [SOP §7.1]
→ Action: suspend Dax + notify Country Safety in 30 min [SOP §2.4]
✅ Severity: P1 true Safety

✓ Every step visible. SOP citations. Conclusion backed by criteria. QA-defensible.

🗳️ Self-Consistency: 3 Approaches, 1 Answer

For high-stakes Safety / fraud / Dax-suspension decisions, analyse from 3 independent angles and take the majority vote:

Approach	Analysis	Conclusion
1. Pax language	Explicit "smelled like beer" — impairment keyword per SOP §8.3	🔴 P1 SAFETY
2. Timing	Ride ended 22:14 SGT, case opened 22:38 — within 1-hour P1 window	🔴 P1 SAFETY
3. Pax credibility	Grab VIP, no past complaints, consistent ride history	🟢 LEGITIMATE REPORT

Majority: 2/3 P1 SAFETY. Approach 3 alone (looking only at Pax credibility) might have de-prioritised this case. The majority vote catches what a single lens misses — and protects the 30-min SLA.

⚠️

GS rule: Any decision that could be audited should use CoT — the reasoning trail IS your documentation.

👤 Role & Persona Prompting

Same data, dramatically different insights — just by changing who the AI "is." The AI was trained on millions of documents written by different professionals. When you assign a persona, you activate that specific knowledge cluster.

The Persona Formula

      You are [TITLE] at [COMPANY TYPE]

      with [X years] of experience in [SPECIALTY].

      You are known for [CHARACTERISTIC].

      When [SITUATION], you always [BEHAVIOR].

💡

The last two fields — CHARACTERISTIC and BEHAVIOR — matter most. "Cautious" vs "opportunity-focused" produces completely different recommendations from the same data.

👀 Same Case, Different Eyes

Data: Pax (Grab VIP) reported "the driver smelled like beer", ride ended 22:14, Dax has a 4.2 rating with one prior unresolved Pax safety report

🛡️ IRT Safety Analyst (default lens)

SEVERITY: P1 — IMMEDIATE ACTION

Explicit impairment language is an automatic P1 trigger per SOP §8.3. Dax has a prior unresolved Pax safety flag — pattern risk.

Action: Suspend Dax pending review. Notify Country Safety in <30 min. First-response email to Pax within SLA.

📞 Pax Care Lead (different lens)

P1 SAFETY + PAX RECOVERY OPPORTUNITY

Pax is a Grab VIP — DSAT and brand risk if not handled with care. Beyond the SOP-required actions, the relationship matters.

Action: Standard P1 escalation + VIP-tier care script. Personalised first response within 15 min, offer goodwill credit, dedicated TL follow-up.

Both are valid. The Safety analyst sees the regulatory + Dax-action picture. The Pax Care lead sees the Pax-relationship picture. Neither is wrong — they serve different audiences (Country Safety lead vs DSAT trend report).

🤝 Multi-Agent Framing

Get 3 perspectives in one prompt — no need to schedule 3 meetings:

Perspective	Focus	Key finding
🛡️ Risk Manager	Default rate, exposure, regulation	"Doubling limits increases exposure by $12M"
📊 Product Manager	Adoption, competition, revenue	"Current $500 limit is #1 reason for churn"
⚖️ Compliance	Responsible lending, MAS guidelines	"MAS requires affordability assessment above $500"

Synthesis: Proceed with phased rollout ($750 first) with income verification. Monitor default rate weekly. Full $1,000 after 90-day review.

🔍

The synthesis is where the real insight lives. No single perspective dominates — the balanced recommendation is stronger than any individual view.

The Research Behind It

Multi-agent framing is a well-established prompt engineering technique with several names in the research literature:

Technique	Source	Key idea
Solo Performance Prompting (SPP)	Wang et al., 2023	A single LLM simulates multiple personas that collaborate internally — "cognitive synergy through multi-persona self-collaboration"
Multi-Persona Thinking (MPT)	arXiv 2025	Dialectical reasoning from multiple perspectives to reduce bias and improve decision quality
Town Hall Debate Prompting	arXiv 2025	Splices a language model into multiple personas that debate one another to reach a conclusion
Self-Consistency	Wang et al., 2022	Generate multiple reasoning paths and aggregate — the broader technique family that multi-perspective builds on

💡

Why it works: LLMs are trained on millions of documents written by different professionals. When you assign a persona, you activate that specific knowledge cluster. Asking for 3 personas in one prompt triggers 3 distinct "knowledge activations" — producing genuinely different analyses, not just rephrased versions of the same answer.

🔗

Day 2 connection: On Day 1, you simulate multiple perspectives in a single prompt. On Day 2, you'll see the agentic version — the Parallelization pattern — where each perspective actually runs as a separate AI agent simultaneously, and a real aggregator combines the results. Same concept, automated at scale.

📋 Structured Outputs & RAG Grounding

Consistent format + grounded in YOUR data = production-safe outputs.

Why Structure Matters

❌ Unstructured = Conversation

Different every time. Hard to compare. Can't feed into systems. Requires human parsing.

✅ Structured = Form

Consistent format. Comparable across items. Machine-parseable. Scannable by busy stakeholders.

How to Prompt for Structured Output

Tell the AI exactly what shape the output should take. The more specific your format instructions, the more consistent the results.

Technique	Prompt example	What you get
Named sections	"Use these sections: Symptom, Severity, Booking, Action, Next Step"	Same headings every case — scannable for TLs and stakeholders
Table format	"Present as: Field \| Value \| Source \| Confidence"	Aligned data, scannable in D365 case notes
JSON output	"Return JSON: {symptom, severity, booking_id, action, sop_citation}"	Machine-readable, feeds into D365 / dashboards / Slack escalations
Numbered actions	"List 3 next agent actions. Each: action, owner (agent / TL / SPV), deadline, SOP-ID"	Actionable items with accountability + SOP traceability
Severity + justification	"Assign P1 / P2 / P3 severity. Justify in exactly 2 sentences with SOP citation."	Consistent QA-defensible severity decisions across all cases
Length control	"Stakeholder summary: max 3 sentences. Detail: max 150 words."	Right depth for the audience (stakeholder vs TL vs agent)

Full Example: Combining Techniques

OUTPUT FORMAT:

1. Risk Rating — GREEN/AMBER/RED with 2-sentence justification
2. Key Metrics Table:
   | Metric | Value | Benchmark | Assessment |
3. Analysis — max 150 words, cite specific numbers
4. Recommended Actions:
   - Numbered list, each with: action, owner, deadline
5. JSON Summary (for system integration):
   {"rating": "...", "confidence": 0-100, "top_risk": "..."}
    

💡

Pro tip: You can mix human-readable sections (1-4) with machine-readable JSON (5) in the same prompt. The AI handles both formats in one response. This is how production templates work — the human reads the narrative, the system reads the JSON.

The Best Default Format: Markdown (.md)

When you ask AI to produce a report, analysis, or any reusable document — ask for Markdown. It's the format that works best for both humans and AI.

Format	Human readable	AI readable	Token cost	Reusable
PDF	✅	❌ Can't parse	N/A	❌
Word (.docx)	✅	⚠️ Partial	N/A	❌
HTML	⚠️ Tags clutter	✅	High (~20 tokens/heading)	✅
Markdown ✓	✅	✅	Low (~8 tokens/heading)	✅

How to ask for it:

Save the output as "case-summary-{case_id}.md" with:
- ## headings for each section
- | tables | for data comparisons
- - bullet lists for action items
    

🔗

Why this matters for you: In today's exercises, every output file is .md. On Day 2, every artifact you create — steering files (.kiro/steering/rules.md), skills (SKILL.md), agent configs — is Markdown. It's the interface layer between you and AI: structured enough for machines, readable enough for humans, and 60% fewer tokens than HTML.

The Research Behind Markdown for AI

This isn't just a convention — research and industry practice back it up:

Finding	Impact	Source
Markdown vs HTML token usage	60% fewer tokens for same content structure	Token comparison (heading: ~8 vs ~20 tokens)
Markdown vs JSON for LLM comprehension	16% average token savings with equal or better accuracy	Format performance benchmarks
Table extraction accuracy	Markdown 60.7% vs HTML 53.6%	ReleasePad, 2025
RAG retrieval with clean Markdown	Up to 35% better retrieval accuracy, 20-30% fewer tokens	AnythingMD
llms.txt web standard (Sept 2024)	Websites now serve Markdown specifically for AI agents	Jeremy Howard, Answer.AI
LLM Markdown awareness research	LLMs are expected to produce structured Markdown for readability	arXiv:2501.15000, 2025

💡

The industry is converging on Markdown as the standard interface between humans and AI. LLMs are trained on it, tools expect it, and it costs less. The llms.txt standard (proposed by Jeremy Howard of fast.ai in September 2024) is like robots.txt but for AI — websites now serve Markdown files at their root specifically for AI agents to read. When you write a steering file, a SKILL.md, or ask for a report — Markdown is the right default.

🔧 Advanced: XML Tags for Claude (Optional)

This section is for Citizen Developers and technical team members. Most business users can skip this — the plain-text techniques above are all you need for daily use.

When building prompt templates at the code level (Bedrock API, application backends), developers often wrap prompt sections in XML tags. This is how Anthropic recommends structuring complex API calls — the tags create unambiguous boundaries between instructions, data, and constraints.

// Typically constructed in application code, not typed by hand:

<role>Senior IRT Team Lead at AnyCompany Support, 8 years handling Safety cases in SEA</role>

<data>
Case ID: BK-2026-4821
Market: SG · Channel: Live Chat · Type: IRT
Pax: P-99421 (Grab VIP, no past complaints)
Dax: D-7711 (rating 4.2, 1 prior unresolved Pax safety flag)
Booking: GR-9821, ride ended 22:14 SGT
Pax message: "the driver smelled like beer when I got in"
</data>

<task>Summarise the case for stakeholder escalation. Assign severity P1 / P2 / P3 with SOP citation.</task>

<constraints>
- ONLY use data in <data> tags
- Max 300 words
</constraints>
    

Pattern	Tags	When to use
Data analysis	`<data>` `<task>` `<format>`	Analyzing reports, transactions, metrics
Document Q&A	`<document>` `<question>` `<rules>`	Policy lookups, compliance checks
Multi-step	`<context>` `<step1>` `<step2>`	Complex workflows, chained analysis
Review	`<draft>` `<criteria>` `<instructions>`	Reviewing reports, emails, proposals

🔍

Where XML tags live in practice: In production systems, the developer builds the prompt template with XML tags programmatically. The end user fills in a form or pastes data — the application wraps it in <data>...</data> tags behind the scenes before sending to the Bedrock API. You design the content; your tech team handles the XML structure. For daily use in Claude Cowork or Kiro, plain-text headers (### ROLE, ### CONTEXT) work just as well.

🔒 RAG Grounding: Only YOUR Data

Without grounding rules, the AI mixes its training data with your documents — you can't trace which claims come from where. In an audit, that's a disaster.

CRITICAL GROUNDING RULES:
Base your answer ONLY on the provided documents
After each claim, cite: [Doc Name, Section]
If not in documents: "Not available in provided documents"
Do NOT use outside knowledge
    

❌ Without RAG Grounding

Q: "What happens if a customer misses 3 PayLater payments?"

The account is typically suspended and sent to collections. Late fees of around 1-2% per month are usually applied.

⚠️ Sounds plausible but NONE of this may be in your actual policy. "Typically" and "usually" = hallucination.

✅ With RAG Grounding

Q: Same question + policy doc attached + grounding rules

Per PayLater Terms v3.2:
• After 3 missed payments → "Default" status [Section 8.3]
• Late fee: $5 or 1.5% whichever higher [Section 7.1]
• [INSUFFICIENT DATA: credit score reporting not in provided document]

✓ Every claim cites a section. Admits what it doesn't know. No hallucination.

🔧 Interactive Prompt Builder

Toggle techniques on/off to see how the prompt AND the AI's response change. Watch quality improve as you add each technique.

👤 Role 📚 Context 📋 Structure 🔒 Grounding 🧠 CoT 🏷️ XML Tags

📝 Your Prompt 0 words

Loading...

🤖 AI Response

Loading...

💡 What changed: Toggle techniques above to see how the AI response improves.

📊 Quality Score

Completeness

2/5

Data Grounding

1/5

Actionability

1/5

Consistency

2/5

6/20

Needs work — toggle more techniques

🔍 Issues in AI Response

⚠️ 7 Prompt Mistakes Everyone Makes

Recognize these patterns? Fix them with one-line additions to your prompt.

Mistake	Why it hurts	Quick fix
🍳 The Kitchen Sink	Cramming 5 tasks into 1 prompt	One task per prompt, chain results
📄 The Blank Canvas	No examples = AI guesses your format	Show 1-2 examples of desired output
🙈 The Trust Fall	No grounding = confident hallucinations	"ONLY from provided data"
🔍 The Vague Ask	"Analyze this" — analyze what, how, for whom?	Specify audience, format, length
⏱️ The One-Shot Wonder	Expecting perfection on first try	Plan for 2-3 refinement turns
📋 The Copy-Paste Trap	Same prompt for different models	Tune syntax per model family
⚙️ The Set-and-Forget	Never re-testing after model updates	Monthly prompt health checks

🔄 The 3-Round Improvement Workflow

Every production-quality prompt goes through this cycle:

Round	What you do	Result
1. Baseline	Write prompt using 4 pillars. Run 3 times.	See what AI gets right and wrong (~60% quality)
2. Fix failures	Add negative constraints + example of good output. Run 3 more.	Consistency jumps to ~85%
3. Polish	Add self-review step. Tighten format. Test edge cases.	Production-ready at ~95%

💡

Total time: 15-20 minutes to go from first draft to production template. That template then saves hours every week.

🚫 Tell the AI What NOT to Do

Negative constraints prevent common failure modes:

Problem	Add this constraint
AI adds unsolicited opinions	"Do not include personal opinions or speculation"
AI uses data not in your input	"Do not reference any data outside the provided documents"
AI writes too much	"Do not exceed 300 words"
AI hedges everything	"Do not use phrases like 'it depends' or 'generally speaking'"
AI explains obvious things	"Do not explain what PayLater is or how digital wallets work"
AI invents numbers	"If a metric is not in the data, write [DATA NOT AVAILABLE]"

🔍

Source: Claude's prompting best practices recommend telling Claude what to do instead of what not to do for general instructions, but negative constraints are highly effective for preventing specific failure modes — especially in finance where hallucinated numbers are dangerous. Claude Prompting Best Practices →

🧠 The #1 Misconception: "AI Remembers Me"

It doesn't. Each session is completely isolated. The AI has zero memory of previous conversations.

❌ What people think

"It remembers our conversation from last week"
"I should keep this tab open so it doesn't forget"
"My old sessions are giving it context"

✅ How it actually works

Each session starts with zero memory
Old tabs have no effect on new sessions
Closing old sessions is safe — cosmetic, not functional

What persists	What doesn't
✅ Files in your workspace (reports, templates, code)	❌ Chat conversation history
✅ Steering files (.kiro/steering/) — loaded every session	❌ What you said 3 sessions ago
✅ Skills (.kiro/skills/) — activated by keywords	❌ Old tabs or closed sessions
✅ Custom agents (.kiro/agents/) — invoked by name	❌ Your "relationship" with the AI

💡

The mental model: chat is ephemeral, files are permanent. Save important outputs as files. Reference files (not old chats) when you need context in a new session. Steering files and skills ARE the AI's persistent memory — they're loaded automatically into every new session.

The 4 Pillars of Prompt Engineering

🎯 The 4 Pillars of Effective Prompts

1. Clarity

2. Context

3. Role Assignment

4. Output Framing

📚 4 Types of Context

👀 See the Difference

🧠 Chain-of-Thought Reasoning

Zero-Shot CoT

Few-Shot CoT

Step-Back

Self-Consistency

👀 Zero-Shot CoT: Before & After

🗳️ Self-Consistency: 3 Approaches, 1 Answer

👤 Role & Persona Prompting

The Persona Formula

👀 Same Case, Different Eyes

🤝 Multi-Agent Framing

The Research Behind It

📋 Structured Outputs & RAG Grounding

Why Structure Matters

How to Prompt for Structured Output

Full Example: Combining Techniques

The Best Default Format: Markdown (.md)

The Research Behind Markdown for AI

🔧 Advanced: XML Tags for Claude (Optional)

🔒 RAG Grounding: Only YOUR Data

🔧 Interactive Prompt Builder

📊 Quality Score

🔍 Issues in AI Response

⚠️ 7 Prompt Mistakes Everyone Makes

🔄 The 3-Round Improvement Workflow

🚫 Tell the AI What NOT to Do

🧠 The #1 Misconception: "AI Remembers Me"