How a small business or remote office can combine two Windows servers into a highly available Hyper-V cluster without purchasing a separate shared-storage appliance.
Article collection
Azure, Azure Local and infrastructure
Architecture and troubleshooting guidance for Azure, Azure Local, Windows Server clustering and software-defined infrastructure.
Published guidance
A backup job reporting success is a claim about a job. The only evidence that matters is a restore you performed, timed, and validated against what the business actually needs.
Peering is not transitive, which is the fact the whole topology is built around. Knowing why it is chosen also tells you the point at which Virtual WAN replaces it.
The reference architecture assumes a dedicated platform function. Most organisations do not have one. The subset that still earns its keep is smaller than the diagram suggests.
Policy is the only control that stops a permitted action producing an unacceptable resource. Deploying it in the wrong order is how you break a platform team's deployments on a Monday.
A private endpoint is a network interface and a DNS problem. Almost every failure I have seen was the DNS half, resolving the public address from somewhere nobody checked.
Role assignments inherit downward and there is no deny. Where you assign a role matters more than which role you assign, and the subscription limit is closer than you think.
Quorum is a voting problem, and the witness is the vote that breaks ties. Each of the three options fails in a different way, and the right choice depends on what you expect to lose.
Live migration fails in a small number of reproducible ways. The most confusing one is authentication, where migrating from the console works and migrating remotely does not.
An Owner who cannot read a secret, a firewall that looks open, and a private endpoint resolving to the wrong address all produce the same status code. A script that separates them.
The storage engine is largely the same software. The operating model is not. Separating those two facts is what makes the platform decision answerable.
The cluster reports its own health accurately. The difficulty is that a degraded volume, a stuck repair job and a failing drive produce overlapping symptoms and need different responses.
A browser-based console that replaces a dozen MMC snap-ins, and a gateway that can reach every server you own. The second half is why its placement is a security decision.
Disk queue length is the counter everyone quotes and the one that misleads most often. Latency is the number that corresponds to what users experience.
In-place upgrade is supported, faster and carries forward everything — including the accumulated configuration nobody understands. That last part is usually the deciding factor.
An architectural view of what changes across Microsoft platform generations, and the operational questions that still need an answer.
