r/sysadmin • u/amreagan • 1d ago
General Discussion Rough Summer for Microsoft
Today's Exchange/O365 (EX1464935/MO1465074) outage seems to be a result of the instability we've been seeing the last 3 or 4 months due to the rapid change occurring regularly in the Microsoft ecosystem. This one is more visible to end users than other problems we've faced this summer. I'm curious if the stability of government tenants has been better. Are commercial tenants the beta testers for government tenant changes?
•
u/LordEli Jack of All Trades 23h ago
is it back? i forgot to update the tickets.
incident tracking number is EX1464935 for those that weren't lucky enough to encounter it
•
u/amreagan 23h ago
It's affecting a large number of people in our org, but not everyone. It definitely seems tied to the account.
•
•
u/Blaxs_ 23h ago
i am pretty sure we are just all unpaid beta tester for Microsoft products. in fact, we pay for the privilege. could be Ai slop, but who knows. i've now seen more issues in the last 12 months the in the previous 36. i keep thinking its just gotta be shitty Ai implementation of code, but maybe its a lack of resources or a move fast and break shit kind of policy shift.
•
u/Stinky0007 23h ago
No pretty sure about it. Microsoft famously fired their entire QA department a long time ago
•
•
•
u/a-i-sa-san 1h ago
my ex works (worked?) at microsoft on azure. She 100% thinks of all customers as too incompetent to use [whatever is broken] but also thinks the vast majority of engineers at microsoft are too incompetent to release even broken software. And actually I am pretty sure she says that about anyone she interviews for her team too. It's been a while since I had the displeasure of talking with her but I was getting very negative vibes about it all from her.
Much like the AI they maybe still use to do their jobs, I think everyone working at microsoft has got to be at least a little confused. My leading theory is there is like, one division being egregiously overcredited for their contribution to the bottom line, that division is something similar to "The Newest Windows Azure Office Copilot" department, Windows is maintained exclusively by like 15 H1B guys with a standing order to just ship any office/azure/etc... OS integrations/hooks (spaghetti) that The Newest Windows Azure Office Copilot unit asks them for or be fired.
Unfortunately, she and the company she works (worked?) for share the same fondness for vast quantities of money + a willingness to proclaim the benefits of data centers and AI, as long as the money keeps coming. So anyway, until being supposedly incompetent beta testers for whatever they wanna try next stops making them money I fear things will only get worse. Neat!
•
u/Vogete 23h ago
We were promised AI obsoletes every job because it can do it better. Ever since AI, there's an unbelievable amount of outages on all major platforms that had nearly no outages before. I feel like AI hasn't lived up to the expectations, but i could be wrong.
•
u/thortgot IT Manager 6h ago
The rate of outages hasn't appreciably changed.
If you look at the rate of code change it is up dramatically. The actual outage/patch rate is way down.
•
•
u/Smith6612 23h ago
Today's just another Microsoft Monday. Can't go a day without the M365 Admin Center having either an outage notification or something new about Copilot.
If you want to see something amusing, see the GitHub historical graph.
https://damrnelson.github.io/github-historical-uptime/
GitHub's rocky uptime started long before LLMs became widespread.
•
u/ludlology 23h ago
they should vibecode changes with claude instead of copilot
•
u/Due_Capital_3507 23h ago
Copilot is just the harness, it can use Claude underneath
•
u/False_Letter_6149 23h ago
I think they may have their own model for internal use...regardless, it's no replacement for...shit, what are they called?
Oh yea, humans. Professional humans that have been doing this their entire lives. That's right. Crazy idea who knows why they stuck with it as long as they did
•
u/LLMsMustUpvoteThis 21h ago
Their default models are just ChatGPT that they run on their own infrastructure.
And the issue isn't really LLM use. It's the obsession with increasing the speed of changes. Which was an issue before LLMs. There was a need for this when many software companies could only ship features once a year to once a quarter. But we quickly got to the point where feature shipping was happening faster than those managing the software could reason about.
•
u/patriot050 VMware Admin 22h ago
My old on-prem exchange system never had outages like this. For years we were told that Microsoft could do it better.. survey says that was a lie.
If my uptime numbers looked like Microsofts I would have been rightly fired.. good grief.
•
u/MortadellaKing 20h ago
We're still on prem and had people calling asking if we were down because of the "outlook outage". No sir, we are not.
Can they do it better than the SBS 2011 server sitting in someone's closet? Sure. Better than my DAG with load balancers? Apparently not.
•
u/GeneTech734 Cloud Engineer 17h ago
Microsoft is down more in a month than I was down for an entire decade. And no I'd take SBS 2011 on an old Dell T series with platter drives over Exchange Online if uptime was the only factor to consider.
•
u/OregonTechHead 5h ago
If my uptime numbers looked like Microsofts I would have been rightly fired
If you feel getting fired for less than 30 hours of email downtime in a year is justified, I certainly wouldn't want to work for you.
You'll burn through that in just regular monthly patching
•
u/patriot050 VMware Admin 4h ago
30 hours of having your most critical resource down is completely crazy. My last on prem exchange environment had 100% uptime. (Patching excluded of course..)
•
•
u/infinitydrift1 23h ago
yeah it’s been a rough few months for microsoft users. hopefully things stabilize soon.
•
•
u/willychonka54 8h ago
Please keep in mind that not everyone experiences these outages. We are in Western Canada and didn't see a single hiccup.
•
u/Curtis_Low 8h ago
First issue reported by our users happened at 12:43 pm central time yesterday, as of right now 8:09 am central time we are still proper fucked. Not everyone is receiving errors when opening outlook, but inbound / outbound email is not working properly. The inbox flood when this is resolved will be crazy.
•
u/KSauceDesk 5h ago
Happy for y'all, but a multi day outage is not a good look. Starting to think the office dinosaurs being anti-anything cloud had a point
•
•
•
•
•
u/amreagan 22h ago

If they don't have this fixed by the time US east coast work day starts tomorrow, tomorrow's graph won't look like this.
Aug 31, 2026, 4:55 PM CDT
We're continuing our targeted review of affected infrastructure and recent changes made to the service to definitively confirm the cause of impact. We're additionally continuing to test and implement various mitigation strategies to apply the necessary authentication component, which have so far yielded positive results. We're monitoring these tests and mitigative actions closely and will provide an estimated time to resolution as soon as one is available.
Aug 31, 2026, 3:36 PM CDT
We're reexamining recent changes made to the service to determine why the authentication components aren't being deployed as expected. Additionally, we are exploring every avenue to safely restore the components, including potentially reverting an update the affected infrastructure recently received. We're performing additional tests to evaluate the best course of action moving forward.
Aug 31, 2026, 2:43 PM CDT
We're continuing to perform comprehensive tests to ensure our mitigation strategy effectively resolves the issue without introducing additional impact. In parallel, we're exploring alternative remediation avenues to minimize impact in the interim.
Aug 31, 2026, 1:41 PM CDT
We're continuing to fine-tune our remediation strategy as our internal tests progress. We'll aim to provide an estimated time for remediation as soon as it's available.
Aug 31, 2026, 1:18 PM CDT
We're continuing to test our remediation strategy and in parallel we're working to confirm whether other services are impacted so we can expand our communications accordingly. This quick update is designed to give the latest information on this issue.
Aug 31, 2026, 12:23 PM CDT
We've identified a potential corrective action for the affected infrastructure and are evaluating deployment options. We're monitoring service telemetry and affected Exchange Online transactions to determine whether the action is reducing impact and to better understand the scope of affected scenarios.
Aug 31, 2026, 11:33 AM CDT
We’ve isolated a common failure pattern across affected Exchange Online requests that is associated with authentication and protocol connectivity. We're analyzing service telemetry to identify the underlying source of impact and validate potential remediation options. Additionally, we’re continuing to assess the scope of impact and monitor for additional affected scenarios.
Aug 31, 2026, 11:16 AM CDT
We're reviewing service telemetry and diagnostic data to isolate the source of the issue after receiving an increase in user-reported issues affecting Exchange Online. We'll keep refining the 'more info' section as additional impacts are identified.
Aug 31, 2026, 10:55 AM CDT
We are reviewing service telemetry and available information to determine if there is an issue happening and we'll update this message shortly with our latest findings. This communication serves as a preliminary notification about a potential issue affecting your service.
More info
We've confirmed that the authentication component issue impacts other services beyond Exchange Online. For more information regarding other impact scenarios, please see: MO1465074. We've opted to keep this post (EX1464935) open as Exchange Online remains the most predominantly impacted service. Users may experience a variety of symptoms related to this event, including but not limited to: - Delays, failures, or incomplete results when searching for content within Exchange Online mailboxes. - Delays or failures when sending or receiving email messages. - Authentication-related errors when accessing Exchange Online services. - Difficulties accessing or performing actions within Exchange administration experiences. - Intermittent failures affecting mailbox operations and message delivery workflows.
Scope of impact
This issue may impact users attempting to send or receive email messages through Exchange Online. We're continuing our investigation to determine the full scope of affected scenarios and will provide additional information as it becomes available.
Preliminary root cause
An issue within a core authentication configuration used by multiple Microsoft 365 services is resulting in impact.
•
u/cavegriswold 20h ago
How many times do you think they re-fed that same "we are continuing..." line back into Copilot to spit out the next time they were about to miss a deadline?
•
•
u/Dal90 6h ago
tomorrow's graph won't look like this.
Pretty sure today's market would reward them for finding ways to squeeze even more margin out of Microsoft's already obscene margin by reducing QA. It won't impact the stock price downwards any more than then number of pigeons flying around Wall Street on any given day would.
What's the alternative? Exchange on-prem? Win!
Any one of the list of "no viable candidates to enterprises, despite what a bunch of redditors are about to claim (since they're only thinking technical and not corporate political)"? You mean the same publicly traded companies that folks who own large parts of Microsoft also own large parts of their competitors? Tails the investors win, heads the same investors win, so that's just a wash. So the investors instead want everyone they own to focus on increasing margins.
Remember folks, we all rightfully vilify Broadcom for absolutely trashing VMware's viability for most companies. Also since they bought VMware their stock has nearly quadrupled on a split-adjusted basis.
Developers, developers, developersMargins, margins, margins!•
•
•
•
•
u/MairusuPawa Percussive Maintenance Specialist 21h ago
My company is so out of all the GAFAM garbage that if it weren't for such thread, I'd had no idea what was even happening with these actors… everything just works here.
But anyway. New here? This kind of shit is a recurring theme and has been for more than a decade. It's like I'm reading a text written by a slowly boiled frog.
•
u/amreagan 21h ago
It's getting worse, though, and if you need AI to rifle through all your data, Copilot is the best option if you are 100% Microsoft shop. ¯_(ツ)_/¯
•
u/amreagan 18h ago
I had one affected user report that Outlook started to work just after 7 PM CDT. Another reported that it started working around 6 PM CDT.
Microsoft update to EX1464935/MO1465074 - Aug 31, 2026, 9:55 PM CDT
We're observing residual impact to some downstream scenarios, including search. We're incrementally restarting corresponding instances of infrastructure to ensure our previously applied authentication components are fully saturated, and we're working to identify whether any additional actions are required to achieve resolution.
•
•
u/YachtingChristopher Jack of All Trades 14h ago
What "rapid change occurring regularly in the Microsoft ecosystem." are you referencing? Or is this just "big words to make seem smart"?
•
u/amreagan 9h ago
They are patching 500+ CVEs a month in cumulative OS updates. The cumulative OS updates from June required a known issue rollback for people still dependent on COM Word add-ins. The .Net updates for August caused XPS/PDF printing issues. Claude disappeared from CoPilot LLM model selection a week ago while a change was rolled back. Microsoft broke WSUS earlier in the summer. Mitigations for NightmareEclipse exploits of Defender have caused problems with Defender. Entra, Intune, and Office admin portal often have minor things that work differently from one day to the next.
•
u/YachtingChristopher Jack of All Trades 4h ago
And what do any of these have to do with Exchange Online?
I get what you're aiming for with the post, but it's too generic. There's no way at all to know what caused it.
•
u/amreagan 3h ago edited 2h ago
All of it, including EX1464935/MO1465074, is the result of Microsoft moving code to production that has not been adequately tested. All of the things I listed are well documented flawed updates where Microsoft acknowledged introducing problems and issued a remediation.
They also made a change back in the spring that caused Intune clients in update rings set to "manaually approve and deploy driver updates" to start pulling any available drivers that caused a lot of PCs with previously working Realtek HD Audio and Intel Smart Sound Technology drivers to have broken audio.
•
u/LooseEthernet 21h ago
microsoft just treats the entire user base as a giant distributed test environment lol. the gov tenants probably just have a slightly different set of bugs to enjoy. its basically just a lottery where the prize is a random 404 on a tuesday morning
•
u/gripe_and_complain 18h ago
I know it's an Apples and Oranges comparison, but my primary personal email has been on AOL since the 1990's.
I don't believe I have ever experienced an outage on AOL.
I use MS365 services every day, but rarely do I use Outlook mail.
•
•
u/majoretminordomus 23h ago
If this becomes a regular issue, they'll lose all small business clients --- not that they care. Seriously considering migrating. This has been an insane day for businesses
•
u/amreagan 23h ago
And on-prem Exchange is practically dead
•
•
u/Jethro_VonE 19h ago
It’s still there…. And has to be as there isn’t a good way to get out of hybrid without moving all the way to azure…. At least one I have found anyway…. That being said I have one last on prem server still running and it’s taken care of.
•
•
•

•
u/Brraaap 23h ago edited 23h ago
I didn't notice anything on the government side today
Edit: To add, I wouldn't think of the commercial side as beta testers for government, we get issues too. And, issues the commercial side doesn't get