We feel more awkwardness in audio calls than video calls and more awk in video calls than physical meetings because at each medium we're losing important contextual cues that act as lubricants to our conversation.
A gentle "slap on knee" while sitting can signal to the other person that you're ready to leave the conversation.
Starring elsewhere while listening can be a sign of "thinking" or "distraction" depending on how your eyes are moving.
In a text-dominant world, where all of these contextual cues are lacking, we tend to interpret people's messages in the most negative way possible. (Snapchat solves this w/ images, teens abuse emojis to solve this, voice msg are becoming more of a thing)
I do think we can incorporate a large chunk of these contextual cues digitally to make digital interactions smoother! Even without VR.
Visual interactions offer increased bandwidth in communication, but in a work exchange I find that bandwidth useless to harmful compared to, say, the intricate process of playing musical instruments in more than perfect sync.
Video calls are additionally worse, the extra input is almost pure noise and cannot even help you read the room to show if a person is distracted when their "listen attentively", "glance on their watch" and "just do something else entirely" are exactly the same. I find it much easier to distinguish with pure audio and no distractions.
Having to use video for work makes me feel as if I was in webcam business, though I admit it is useful in more sensitive meetings where you want to visually confirm participants.
Audio calls mostly just feel like phone calls to me, except without having to hold a phone up to my ear. Even if I'm using Discord or Slack or some company-specific thing, audio calls trigger "muscle memory" that I've built up my whole life. Video calls are much more novel to me; I rarely ever did it before covid, and although it now is a bit more natural, it'll be a long time until it feels as "normal" as audio calls for me.
There's also the latency. POTS audio added very little, but modern codecs add some and modern networks add some more, and modern computer audio adds a bit more and soon enough it's enough to make the conversation harder, even when it's hard to perceive.
A gentle "slap on knee" while sitting can signal to the other person that you're ready to leave the conversation.
Starring elsewhere while listening can be a sign of "thinking" or "distraction" depending on how your eyes are moving.
In a text-dominant world, where all of these contextual cues are lacking, we tend to interpret people's messages in the most negative way possible. (Snapchat solves this w/ images, teens abuse emojis to solve this, voice msg are becoming more of a thing)
I do think we can incorporate a large chunk of these contextual cues digitally to make digital interactions smoother! Even without VR.