Tickd.ai
← The Tickd Guide

Future of AI

Why You Should Stop Writing Prompts and Start Compiling Your LLM Workflows

Hand-crafted prompts are fragile, non-portable, and impossible to scale. The future of AI engineering belongs to programmatic, compiled prompt frameworks like DSPy.

Updated 9/17/2026

The Fragility of the English Compiler

We have been told a beautiful lie: that the natural language prompt is the ultimate programming language of the AI era. We have spent countless hours carefully massaging adjectives, adding desperate pleas like "take a deep breath and think step-by-step", and wrapping our inputs in elaborate XML tags to keep our LLMs on the straight and narrow.

But anyone who has tried to move a complex system from a prototype to a production-grade application knows the painful truth. Prompt engineering is a house of cards.

You spend weeks perfecting a prompt for Claude, only to find that it fails spectacularly when run on Gemini. Or you update your system prompt by changing a single word, and the model's output formatting completely falls apart. It is a manual, unscientific process of trial and error that looks less like software engineering and more like medieval alchemy.

It is time to stop writing prompts by hand. The future of AI development belongs to programmatic, compiled workflows.

The Problem with Manual Prompting

When we write static prompts, we are hardcoding behaviour. This breaks several core principles of robust software development:

  • Zero Portability: A prompt optimised for GPT-4o will not perform reliably on an open-source model like Llama-3. You have to start the manual tuning process all over again. You can see how these discrepancies play out in our guide to Claude vs Gemini, where slight structural differences radically alter response quality.
  • Brittle Pipelines: If your application relies on chaining multiple LLM steps together, a slight shift in the output format of step one can cause a catastrophic failure in step three.
  • No Optimisation Loop: How do you know if your prompt is actually the best version possible? Without a systematic way to test, score, and iterate on prompt variations, you are just guessing in the dark.

If you are still manually tweaking system strings inside your API calls, there is a ticking clock on your application. A subtle model update behind the scenes can tick off your entire user base by breaking your carefully crafted outputs. For troubleshooting APIs when these unexpected shifts happen, developers often have to dig through forums like the Gemini Support Site to find out what changed in the underlying architecture.

Enter DSPy: Compiling Instead of Writing

The shift away from manual prompting is being led by frameworks like Stanford’s DSPy (Declarative Self-improving Language Programs). Instead of treating prompt engineering as an art form, DSPy treats it as an optimisation problem.

Instead of writing a massive prompt containing instructions, examples, and formatting rules, you write clean, modular Python code. You define:

  1. Signatures: The inputs and outputs of your system (e.g., question -> answer).
  2. Modules: The structural flow of your reasoning (e.g., ChainOfThought or MultiHopRetrieval).
  3. Teleprompters (Optimisers): Algorithms that programmatically generate, test, and select the best prompts and few-shot examples for your specific model based on a training dataset.

This is a paradigm shift. You do not write the prompt; the framework compiles the prompt for you.

If you want to switch your backend model from OpenAI to an open-source model running locally, you don't rewrite your instructions. You simply point DSPy to the new model and run the compiler. The framework automatically discovers the optimal prompting strategy, formatting, and few-shot exemplars that make that specific model perform best for your task.

How It Works in Practice

Let’s compare the traditional approach with the compiled approach.

In the traditional paradigm, if you wanted to build a system that extracts key entities from legal documents and rates their risk, you would write a three-page prompt filled with legal jargon, negative examples, and strict JSON output structures. You would paste this into our prompts generator to get a clean baseline, but you’d still be stuck manually maintaining it as laws or models change.

In the compiled paradigm, you write a simple signature:

`python class LegalRiskExtractor(dspy.Signature): """Extract legal entities and evaluate compliance risk levels.""" document = dspy.InputField(desc="Raw legal contract text") entities = dspy.OutputField(desc="List of identified entities and their risk scores") `

You then provide a small dataset of 20 to 50 examples of your expected inputs and outputs. You define a metric—say, a validator function that checks if the output is valid JSON and contains the correct keys.

When you compile the program, DSPy runs an evolutionary search. It tests different ways of phrasing the instructions, bootstraps the best examples from your dataset to use as few-shot context, and evaluates the performance. The output is a highly optimised, model-specific prompt that guarantees far higher accuracy and reliability than anything a human could have written by hand.

The Future Is Programmatic

We are moving towards a world of compound AI systems—complex pipelines where multiple small, highly specialised models talk to each other, fetch external data, and call APIs. Trying to manage these networks with hand-written prompts is like trying to build a modern database using only basic spreadsheet formulas. It does not scale.

By treating prompts as compiled artifacts rather than static source code, we unlock the true potential of LLMs. We can finally version-control our logic, run automated regression tests on our AI pipelines, and swap models in and out of our stack without fear of breaking our entire application.

Stop wasting your time guessing which adjectives your LLM prefers. Let the compilers do the dirty work.

future-of-aiprompt-engineeringdspysoftware-engineeringllmops

Keep going

Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.