How to Enforce Team Coding Standards with AI
Macroscope
Macroscope
Product

How to Enforce Team Coding Standards with AI

A practical guide to enforcing team coding standards with AI: what belongs in a linter, CODEOWNERS, or an AI check, how to write custom AI code review rules in plain-English Markdown that run as GitHub check runs and block merges, example check definitions, and how to roll them out.

The best way to enforce team coding standards with AI is to keep deterministic rules in linters and formatters, keep ownership in CODEOWNERS, and move the standards that need judgment into AI checks: plain-English rules that run on every pull request as GitHub check runs and can block a merge when they fail. Each layer does what it is good at, and AI checks cover the standards that today live only in senior engineers' heads.

Every team has rules that never made it into tooling. "Don't query the database from the frontend." "Every migration must be reversible." These get enforced by whoever happens to review the PR, which means inconsistently, and the gap widens as coding agents write more of the code.

This guide covers which standards belong in which tool, how to write custom AI code review rules as Markdown with Macroscope's Check Run Agents, four example check definitions, and how to roll checks out from advisory to blocking.

TL;DR: Enforcing coding standards with AI

  • Use three layers. Linters and formatters for deterministic rules, CODEOWNERS for who must review, AI checks for standards that need judgment or context.
  • AI checks are for rules a linter can't express: architectural boundaries, migration safety, analytics event integrity, security and compliance conventions, docs and SEO metadata.
  • Write each rule like you would explain it to a new senior engineer. With Check Run Agents, each rule is a Markdown file that runs as its own GitHub check run.
  • Checks can block merges. Set the conclusion to failure and require the check in branch protection.
  • Roll out advisory first, tune the wording on real PRs, then make the important ones blocking.

Why are team coding standards so hard to enforce?

Most team standards require judgment, and the only judgment in the pipeline today is a human reviewer's. Linters enforce what can be pattern-matched. Everything else depends on the right person reviewing the right PR on the right day.

  • Tribal knowledge does not scale. The rules that matter most are rarely written anywhere a tool can read.
  • Review attention is limited. Under a heavy queue, reviewers focus on correctness and skim conventions.
  • Agents follow written rules only. More agent PRs means more drift from conventions nobody encoded.

The fix is not a longer style guide. It is turning the standards your team cares about into checks that run on every GitHub PR review.

Linters vs CODEOWNERS vs AI checks: what should enforce what?

Use linters and formatters for anything deterministic, CODEOWNERS for who must approve, and AI checks for anything that needs context or judgment. They are complements. A good setup uses all of them.

LayerBest atExamplesCan't do
Formatters (Prettier, gofmt, Black)Consistent style, zero debateWhitespace, quotes, line lengthAnything about meaning
Linters and static analysis (ESLint, Semgrep, golangci-lint)Fast, deterministic pattern rulesNo eval(), import order, banned functionsRules that depend on intent or other files
CODEOWNERSRouting review to the right people/billing needs a payments engineerJudging the change itself
Required CI testsVerifying behaviorUnit, integration, type checksConventions tests don't encode
AI checks (Check Run Agents)Standards needing judgment or cross-file and cross-system contextArchitecture boundaries, migration safety, analytics events, auth on new routesReplace fast deterministic tools cheaply
Human reviewDesign, product intent, tradeoffs"Is this the right approach?"Scale to every PR

The rule of thumb: if you can write it as a regex or an AST pattern, put it in a linter. It will be faster, cheaper and perfectly consistent. If explaining the rule needs the word "unless", or information from another file, another system or the PR description, it belongs in an AI check.

Which rules belong in an AI check instead of a linter?

A rule belongs in an AI check when enforcing it requires understanding what the change does, not just what it looks like. Five categories come up on almost every team.

  • Architectural conventions. "The frontend never queries the database directly." "Package A must not import package B's internals." "Background work goes through the queue." Easy to say in English, painful as lint rules.
  • Migration safety. Is the migration reversible? Does it drop a column that application code still reads? Does it add a non-null column without a default? A linter sees SQL. An AI check sees the SQL and the code that depends on it.
  • Analytics event integrity. Tracking breaks quietly when someone renames an event or drops a call in a refactor. A check scoped to analytics calls can flag every removed or renamed event before merge.
  • Security and compliance conventions. Your codebase's own rules: new routes use the auth middleware, sensitive paths write an audit log, PII goes through the redaction helper. Keep dedicated scanners for known vulnerability patterns.
  • Docs and SEO metadata. Every page needs a title, description and social image; image references must exist; public API changes need a changelog entry. Context-heavy, easy to forget.

How do you write custom AI code review rules in Markdown?

With Macroscope's Check Run Agents, each custom AI check is one Markdown file in .macroscope/check-run-agents/ on your default branch: optional YAML frontmatter, then plain-English instructions. Macroscope runs each agent on every matching pull request as its own GitHub check run, alongside your CI. Full reference: Check Run Agents: custom AI checks for pull requests.

The main frontmatter fields:

  • title: the check name in GitHub's Checks tab.
  • include / exclude: globs that scope the check. If a PR touches no included files, the agent doesn't run.
  • tools: by default the agent can browse code, read git history, read GitHub metadata and comment on the PR. Optional tools include Sentry, PostHog, Amplitude, LaunchDarkly, BigQuery, GCP logs, Jira and Linear, Slack, web search and any MCP server.
  • input: incremental (default) reviews files changed since the last review; full_diff gives one agent the whole PR; code_object runs up to 20 agents in parallel, one per changed function or class; pr_metadata reads only title, description, labels and commits.
  • model, effort, reasoning: which model runs and how deeply it investigates.
  • conclusion: neutral (default) reports without blocking; failure makes the check a merge gate. Blocking requires input: full_diff.

Findings appear in the Checks tab, as inline comments on specific lines, and as top-level PR comments.

Writing instructions that work: be specific ("flag any new route under src/api/ that does not use requireAuth", not "check security"), say what to ignore (pre-existing issues, tests, generated code), define what counts as blocking, and explicitly allow the agent to report that nothing was found. Don't duplicate the correctness review Macroscope's AI code review already runs; custom checks should encode your team's standards.

What do example AI check definitions look like?

Below are four illustrative check definitions. They are examples to adapt, not drop-in files; paths and helper names are placeholders for your own.

Example 1: Architecture boundaries

---
title: Architecture Boundaries
include:
  - "web/**"
  - "services/**"
---

Enforce our service boundaries on this PR:

1. Code under web/ must never import database clients or run SQL.
   Frontend data access goes through the API service.
2. Code under services/billing/ must not import services/accounts/
   internals. Use services/accounts/client/ instead.
3. New background work must go through jobs.Enqueue.

Only flag violations this PR introduces. Comment inline on each
violating line with the rule it breaks and the approved alternative.
If there are no violations, say so.

Example 2: Migration safety (blocking)

---
title: Migration Safety
input: full_diff
conclusion: failure
effort: high
include:
  - "db/migrations/**"
  - "app/models/**"
---

Review any database migration in this PR:

- Every migration must have a working down migration.
- Dropping or renaming a column is blocking if application code
  still reads or writes it.
- Adding a NOT NULL column without a default is blocking.

Fail the check only for the blocking cases. Report anything else as
a suggestion. If there are no migrations, report that the check passed.

Example 3: Analytics event protection

---
title: Analytics Events
tools:
  - browse_code
  - git_tools
  - modify_pr
  - slack
include:
  - "src/**/*.ts"
  - "src/**/*.tsx"
---

Look for changes to analytics tracking calls (track, capture, logEvent).

For each event removed, renamed, or with changed properties, list the
old and new form, add the label "analytics-change" to the PR, and post
a short summary with the PR link to #data-team in Slack.

If no tracking calls changed, report that the check passed.

Example 4: Docs and SEO metadata

---
title: Content Metadata
input: full_diff
include:
  - "content/**"
---

For every new or modified page under content/:

- Frontmatter must include title, description, and socialImage.
- Every image path referenced must exist in the repository.
- Internal links must point to pages that exist.

List each problem with the file and the fix. If everything is in
order, say so.

Each of these is a rule a senior reviewer would catch by reading the diff and checking some context. None is practical as a lint rule.

How do you roll out AI checks without slowing the team down?

Start every check as advisory, tune it on real PRs, then make the high-value ones blocking. A blocking check that fires on false positives teaches engineers to ignore it.

  1. Mine your review history. Pydantic pointed an agent at a year of its human PR comments, extracted what its senior engineers flagged most often, and turned those into Check Run Agents, per the Pydantic case study. Your repeated review comments are your real standards.
  2. Ship two or three checks as advisory. Leave conclusion at neutral, so findings appear but nothing blocks.
  3. Tune for a week or two. Every false positive is a wording fix: add ignore clauses, tighten scopes, define severity.
  4. Promote critical checks to blocking. Set conclusion: failure with input: full_diff, then require the check in GitHub branch protection. Reserve this for rules where a violation should genuinely stop a merge.
  5. Scope per area. Macroscope runs six style-guide checks internally, one per language, each scoped to its own paths (Introducing Check Run Agents). Narrow scope means less noise and lower cost.
  6. Treat checks like code. They live in the repo, so changing a standard goes through review and has history.

Encoded standards also make AI auto-approval safer. See how to speed up pull request approvals safely.

What is agentic CI?

Agentic CI is continuous integration where checks investigate and reason about a change instead of only running a fixed script. Traditional CI is procedural: run this command, pass or fail. Agentic CI is closer to delegating a review task: here is the standard and the tools, determine whether this PR meets it.

One agentic check can read the linked Jira or Linear ticket to confirm the PR does what was specified, query Sentry for errors in modified files, and post to Slack, all inside one GitHub check run. Check Run Agents are Macroscope's implementation of agentic CI for pull requests. Tests and linters still run; agentic checks cover the layer in between.

How do AI checks compare across tools?

Most AI code review tools let you customize review behavior; fewer let you define independent checks that run as GitHub check runs, use external tools, and block merges.

Macroscope Check Run AgentsCodeRabbit custom rulesGreptileCustom GitHub Actions
Defined asMarkdown file per check.coderabbit.yaml configurationCustom rules plus learning from PR commentsCode and YAML you maintain
Own GitHub check runYesPart of the reviewPart of the reviewYes
External tools (Sentry, Jira, Slack, MCP)Built inNot from custom rulesNot from custom rulesBuild it yourself
Can block mergesYes (conclusion: failure)YesVia required reviews in branch protectionYes
MaintenanceEdit MarkdownEdit YAMLLow, learned over timeCode, prompts, keys, infra

To be fair: Greptile's learned approach needs little setup, at the cost of explicit control. CodeRabbit's path instructions fit teams that mainly want to tune one reviewer by directory. Custom Actions give total control if you will own the infrastructure. See CodeRabbit alternatives and Greptile alternatives.

Two limits, stated plainly: Macroscope is GitHub only today (no GitLab, Bitbucket or self-hosting), and Check Run Agents investigate and report rather than run tests or edit files. For fixes, Fix It For Me opens fix PRs for correctness findings and iterates until CI passes.

What does it cost to run AI checks?

Check Run Agents bill as Agent usage: $0.01 per credit, with 1,000 free Agent credits per workspace per month. There are no per-seat fees, and new workspaces start with $100 in usage credit. Tight include scopes, full_diff instead of code_object, and lower effort on simple checks keep spend down, and monthly budget limits cap it. See pricing.

How do you get started?

Install Macroscope on GitHub, add one Markdown file to .macroscope/check-run-agents/, and merge it to your default branch; it runs on the next PR. The GitHub setup guide covers installation. Start with the standard your reviewers repeat most often.

Need better visibility into your codebase?
Get started with $100 in free usage.

Frequently Asked Questions

What's the best way to enforce team coding standards with AI?

Split standards by type. Keep deterministic rules (formatting, banned functions, import order) in linters and formatters, ownership in CODEOWNERS, and judgment-based standards (architecture boundaries, migration safety, analytics events, team security conventions, docs metadata) in AI checks that run on every pull request. With Macroscope's Check Run Agents, each standard is a plain-English Markdown file that runs as a GitHub check run. Roll checks out as advisory, tune them, then make the important ones blocking.

How do I add custom AI checks to GitHub pull requests?

With Macroscope installed as a GitHub App, add a Markdown file to .macroscope/check-run-agents/ on your default branch. Optional YAML frontmatter sets the title, file scope, tools, input mode and whether it can block; the body is your instructions in plain English. On the next pull request it runs as its own GitHub check run, with results in the Checks tab, inline comments and PR comments.

Can I write custom AI code review rules?

Yes. With Check Run Agents you write each rule in plain English: what to look for, what to ignore, what counts as blocking, and how to format findings. Each file is an independent check, so you can have one for architecture boundaries, one for migrations, one for analytics events. Other tools offer custom rules too, including CodeRabbit through .coderabbit.yaml and Greptile alongside learning from your PR comments.

What AI tool enforces engineering standards on every PR?

Macroscope's Check Run Agents run your custom standards on every matching pull request as GitHub check runs, can use tools like Sentry, PostHog, LaunchDarkly, Jira, Linear and Slack, and can block merges with conclusion: failure. Alongside them, Macroscope's built-in AI code review looks for correctness bugs and its Approvability check decides which PRs are safe to auto-approve. Macroscope is GitHub only today.

What is agentic CI?

Agentic CI is continuous integration where checks investigate and reason instead of only running a fixed script. An agentic check gets a standard in natural language plus tools (codebase browsing, git history, issue trackers, observability and analytics) and determines whether a pull request meets it. It complements deterministic CI like tests and linters. Check Run Agents are Macroscope's implementation of agentic CI for GitHub pull requests.

Can an AI check block a pull request from merging?

Yes. Set conclusion: failure with input: full_diff in the check's frontmatter, then mark the check as required in GitHub branch protection. When the agent finds a blocking violation, the check fails and the PR cannot merge until it is fixed. The default, neutral, reports without blocking.

How much do custom AI checks cost?

Check Run Agents bill as Agent usage at $0.01 per credit, and every workspace gets 1,000 free Agent credits per month. There are no per-seat fees, and new workspaces start with $100 in usage credit. See pricing.

Do Check Run Agents work with GitLab or Bitbucket?

No. Macroscope supports GitHub only today, through its GitHub App, and there is no self-hosted option.

Need better visibility into your codebase?
Get started with $100 in free usage.