Use AI to explore evidence and accelerate a focused investigation. Start with a measurable problem and approved diagnostic material, then ask an agent to rank hypotheses, identify missing evidence, and propose the smallest experiment. A performance expert owns the interpretation and the release decision. An attractive explanation or a faster local run does not establish that the customer problem is solved.
Establish the performance contract
Choose one customer operation and define the load under which it must work. Record the application commit, runtime, instance size, database state, workload mix, data volume, concurrency or arrival rate, warm-up, and measurement window. Agree the latency percentiles, throughput, error rate, and resource limits that matter before making changes.
- Reproduce the symptom with representative test data and record repeated baseline runs.
- Measure achieved throughput alongside latency; lower latency caused by completing less work is not the same result.
- Separate startup and cache-warm behavior from steady-state measurements when both matter.
- Keep the load generator and dependencies observable so their saturation is not mistaken for application cost.
Ask the agent to investigate a specific question
Provide a sanitized evidence summary and only the code paths needed for the investigation. Require every hypothesis to name its supporting observation, a competing explanation, and a discriminating check. The agent may propose a diagnostic command or a code change; an engineer chooses whether and where to run it.
- Use runtime counters for first-level signals and collect a targeted trace when a specific question needs deeper evidence.
- Distinguish CPU work, blocked threads, allocation and collection pressure, dependency waits, and database work.
- For suspected ThreadPool starvation, correlate queueing and worker behavior with stacks or traces; low CPU alone is not a diagnosis.
- For suspected query cost, inspect generated SQL, round trips, result size, and the relevant execution plan before choosing a remedy.
- Do not upload production dumps, credentials, or unreviewed customer data to an AI service.
Change one mechanism and preserve the business result
Choose the smallest fix supported by the evidence and state what observation should change if the hypothesis is correct. For example, eliminate a measured repeated query, remove a demonstrated blocking call, or reduce an identified allocation hot path. Keep unrelated refactoring outside the experiment.
- Add or retain functional checks for authorization, ordering, totals, errors, and other behavior the change could affect.
- Review query and index changes for write cost, contention, and representative parameter values.
- Treat caching as a design change with explicit freshness and invalidation rules, not a default performance patch.
- Require a human review of the diff, the diagnostic reasoning, and any assumptions the agent could not verify.
Verify the change under the same conditions
Repeat the original workload with the same measurement method and report both the improvement and any tradeoff. Keep raw results and configuration with the change so another engineer can reproduce the comparison. After a controlled rollout, compare production signals with the agreed stop conditions.
- Compare latency distributions, achieved throughput, errors, CPU, memory, and relevant dependency measurements.
- Use sufficient run duration to expose the original problem; a short run cannot rule out a slow memory-growth issue.
- Explain normal run-to-run variation and avoid presenting a single best run as the result.
- Keep a rehearsed recovery path and record any remaining performance limits.
Illustrative agent brief: an intermittently slow order-history endpoint
Investigate this endpoint using the approved trace summary, generated SQL, and bounded source files. Return a ranked hypothesis table with supporting evidence, alternative explanations, and the next discriminating measurement. Do not change production settings, add caching, remove authorization, or rewrite unrelated code. Propose one bounded fix only after the evidence supports it.
Acceptance criteria
- The baseline and changed build run the same documented workload against comparable data and infrastructure.
- A trace, query observation, or other reproducible measurement supports the selected cause.
- The agreed latency, throughput, and error thresholds are evaluated together; no percentage gain is asserted without its measurement context.
- Functional and authorization checks pass, and relevant resource or database regressions are explicitly reviewed.
- A reviewer can reproduce the comparison and identify the rollout stop condition and recovery procedure.
The artifacts to keep
- Representative workload and baseline
- Hypothesis and evidence log
- Reviewed bounded fix
- Before-and-after measurement pack
- Regression checks and rollout notes
Limits to keep visible
- These are investigation steps and illustrative acceptance criteria, not a benchmark result or a promised performance gain.
- Diagnostic tool support depends on the runtime and operating system. The linked dotnet-counters and dotnet-trace guidance concerns modern .NET; choose compatible tooling for .NET Framework applications.
- An AI assistant operates only on the evidence supplied to it. A confident explanation does not verify causation, and profiling itself can affect observed performance.
Use the agent-ready kit to scope an investigation, then bring the baseline and evidence gaps to a performance review.
Primary sources
- Microsoft: dotnet-countersFirst-level performance investigation and runtime counter collection.
- Microsoft: dotnet-traceTargeted diagnostic trace collection for modern .NET.
- Microsoft: Debug ThreadPool starvationInvestigating ThreadPool starvation with runtime signals and diagnostic evidence.
- Microsoft: Efficient querying in EF CoreQuery plans, indexing, result sizes, projections, and query round trips.
Co-written by Adnan Rafiq and ChatGPT. ChatGPT assists with research, drafting, and source checks; the engineering direction comes from Adnan. Examples are illustrative unless an article explicitly provides executed results. These are working methods, not customer case studies or guaranteed outcomes.