Probably a Bad Idea...
Posted: Fri Feb 07, 2025 10:22 pm
I've had this idea for a while now that my lack of understanding, as to what actually happens that allows a sample to accurately represent a moment in sonic time, has prevented me from getting my arms around. I keep planning to educate myself on the subject - but... then I don't. I think I know that each sample represents the complex wave at the moment of the sample. But - how does it represent it? Is it just a very particular amount of pressure in that moment? Like, is 16 bit 65,536 different, step-wise, variations in air pressure between silence and the limit of the system's head room? And 24 bit, 16,777,216 variations?
This quandary is all in service to the idea I had. I was reading about how the particular magic of the new 50-series of NVidia graphics cards is about rendering frames that don't actually exist in the graphic recording by, essentially, looking at the value of two consecutive frames for each single dot on the screen and coming up with an average value between them that might well represent what that dot would be, had it been recorded at a higher frame rate. So, presuming grey scale, for ease in example, if a dot had a blackness of 4 in the first frame, and a blackness of 6 in the second - the dot for the generated frame between them would be 5 - half way in between 4 and 6. It's called 'interpolation'. It's a pretty slick idea, I think. And they are stuffing more than one frame in between every other actual frame - resulting in really high apparent frame rates without the burden of having to store them or send them over the wire.
And I wondered why that couldn't happen in sound recording as well. Like, in between each recorded sample at 44.1 sps, you would add, in real time, one that would represent a pretty good guess at what the wave might well have looked like at that moment in time (that was not recorded) if it had been recorded - resulting in double the resolution.
Mind you, I'm not asking whether it's a good idea or not - like, ethically or morally (emphatically intending to head off that train...). Just whether it is possible using however a sample is represented, and whether the resulting synthetic samples are likely to be sort of accurate in representing the complex wave at that (unrecorded) moment in time?
My reasoning for believing it might be is that a sound wave, by definition, is changes in relative air pressure, which is an analog phenomena, and continuous, and must be all possible values in pressure in between those pressures represented by the two samples at some point in time between them, as the pressure either grows or recedes. At the half-way point in time between them - wouldn't the pressure likely be half way between too? If not - would it necessarily be at least very close?
I said it would happen 'in real time'. Actually, I guess it wouldn't necessarily have to (it would be a nifty job for an RTX 5090 though). I suppose you could use less compute to convert a lower resolution file into a higher resolution file, save it - and then play it back too. That way it could take longer to convert than it would take to play. But, of course, then your 5090 is just sitting around chewing gum...
I envision a stereo with a button you could press to cause the music you're listening to to be interpolated to twice as many samples per second in real time, resulting in a change in 'feel' that some might like.
Another possible use would be in 'compressing' sound files for digital transport. The 'decompression' on the other side would just be interpolating samples in between the ones sent.
Anyway, what do you think? Anyone with mad digital recording knowledge care to comment? Is this feasible? I'm hoping Naturally Digital is still around and can weigh in. And Bob if he's not busy and it makes his propeller spin. But any others with knowledge are also encouraged. And - if it's a bad idea technically for some reason I haven't thought of - please educate my too-curious lazy ass as to why - and then I can stop thinking about it.
This quandary is all in service to the idea I had. I was reading about how the particular magic of the new 50-series of NVidia graphics cards is about rendering frames that don't actually exist in the graphic recording by, essentially, looking at the value of two consecutive frames for each single dot on the screen and coming up with an average value between them that might well represent what that dot would be, had it been recorded at a higher frame rate. So, presuming grey scale, for ease in example, if a dot had a blackness of 4 in the first frame, and a blackness of 6 in the second - the dot for the generated frame between them would be 5 - half way in between 4 and 6. It's called 'interpolation'. It's a pretty slick idea, I think. And they are stuffing more than one frame in between every other actual frame - resulting in really high apparent frame rates without the burden of having to store them or send them over the wire.
And I wondered why that couldn't happen in sound recording as well. Like, in between each recorded sample at 44.1 sps, you would add, in real time, one that would represent a pretty good guess at what the wave might well have looked like at that moment in time (that was not recorded) if it had been recorded - resulting in double the resolution.
Mind you, I'm not asking whether it's a good idea or not - like, ethically or morally (emphatically intending to head off that train...). Just whether it is possible using however a sample is represented, and whether the resulting synthetic samples are likely to be sort of accurate in representing the complex wave at that (unrecorded) moment in time?
My reasoning for believing it might be is that a sound wave, by definition, is changes in relative air pressure, which is an analog phenomena, and continuous, and must be all possible values in pressure in between those pressures represented by the two samples at some point in time between them, as the pressure either grows or recedes. At the half-way point in time between them - wouldn't the pressure likely be half way between too? If not - would it necessarily be at least very close?
I said it would happen 'in real time'. Actually, I guess it wouldn't necessarily have to (it would be a nifty job for an RTX 5090 though). I suppose you could use less compute to convert a lower resolution file into a higher resolution file, save it - and then play it back too. That way it could take longer to convert than it would take to play. But, of course, then your 5090 is just sitting around chewing gum...
I envision a stereo with a button you could press to cause the music you're listening to to be interpolated to twice as many samples per second in real time, resulting in a change in 'feel' that some might like.
Another possible use would be in 'compressing' sound files for digital transport. The 'decompression' on the other side would just be interpolating samples in between the ones sent.
Anyway, what do you think? Anyone with mad digital recording knowledge care to comment? Is this feasible? I'm hoping Naturally Digital is still around and can weigh in. And Bob if he's not busy and it makes his propeller spin. But any others with knowledge are also encouraged. And - if it's a bad idea technically for some reason I haven't thought of - please educate my too-curious lazy ass as to why - and then I can stop thinking about it.