I am convinced that there may be a performance bottleneck somewhere in the P2P system of the Hive blockchain software.
The peer to peer network is behaving fine, no bottlenecks there, it could handle much more traffic than current traffic.
I'm currently sat in front of my witness console watching the logs go by and I'm seeing on average over 300 transactions per block.
Yep, but 300 transactions per block is no problem for the network. We could handle far more if we increased the block size (or if there were fewer "big" transactions).
This all seemed to start when a couple of accounts started posting a lot of transactions with nearly depleted RC. It pops up a warning saying they don't have enough RC for the transaction then it's later included in the next block anyways because that's how it works. (there is a grace thing to accept transaction if you would have regenerated the RC by the next block or something)
The accounts you see that are continually creating transactions when they don't have RC are mostly bots that keep retrying because they don't bother to detect they have run out of RC. It is just lazy programming by the bot owners.
If your node is public, they may be pointing directly at you. If they are using broadcast_transaction_synchrounous instead of broadcast_transaction, that's the likely cause of some of your problems. Other nodes won't have the same problem unless they are also pointed to by such bots.
We've asked users to stop using this, but some people didn't get the memo. Your simplest solution if your node is public would be to interpose a proxy and block such bad traffic.
The other thing I'm noticing is Peer churn, where peers I'm currently connected to are not responding fast enough and are being dropped.
It sounds probable that you may be connecting to some steem nodes. They are not completely separated from our network yet, as the network separation is being done in phases over a couple of hardforks.
Another possibility is that your node is being slowed by direct broadcast traffic and is reacting slowly to those nodes.
The final possibility is your disks may be too slow (if you're just using HDDs). Also, hopefully you are using tmpfs to store your shared_memory.bin file.
Along with this I'm seeing block time offset also increase sometimes as high as 3000ms.
Average block time offfset on nodes I've checked here is between 70-300ms. I've seen a few outliers around 500ms.
Is there a performance issue here?
Is it my node? (Highly unlikely)
Your node is definitely suffering a performance problem. It may be due to your traffic or it may be due to your hardware. Hard to say without knowing more about your setup.
Is it because everyone elses nodes are struggling?
None of our nodes are having problems. I've checked several in the cloud and local ones in our data centers. These are all running on good but not great hardware.
Do we just need better hardware now for consensus witnesses? (by consensus I mean normal hive nodes not just the top 20)
All nodes need decent hardware. Any consensus node needs the same hardware as a top 20 node. The real expense in hardware comes if you want to operate a public api node, but even that isn't too expensive really, you just have to be sure you have enough network upload bandwidth with your ISP.
RE: Observation from a witness. I'm I seeing things?